From a t Statistic to a P-Value
In “The One-Sample t Test Statistic” and “Why the t Test Uses n Minus 1 Degrees of Freedom,” you learned how to calculate the observed \(t\) statistic and choose its degrees of freedom. The next step is to find the area under the appropriate t curve that is at least as extreme as the observed statistic in the direction or directions specified by the alternative hypothesis. That area is the p-value.
The calculator function tcdf finds the area under a t curve between two horizontal values. For a one-sample t test, enter the lower bound, upper bound, and \(df=n-1\). On a TI-84, very large positive and negative numbers such as \(1\text{E}99\) and \(-1\text{E}99\) stand in for infinity when asking for an entire tail.
The alternative hypothesis determines which area to calculate. A lower-tailed alternative asks for the area to the left of the observed statistic. An upper-tailed alternative asks for the area to its right. A two-sided alternative counts results at least as far from 0 in either direction, so its p-value includes both tails. These rules use the t curve with the test’s own degrees of freedom.
- For \(H_a:\mu<\mu_0\): area to the left of \(t\).
- For \(H_a:\mu>\mu_0\): area to the right of \(t\).
- For \(H_a:\mu\ne\mu_0\): area beyond \(-|t|\) and \(+|t|\).
On a calculator, those areas can be requested as follows. Use the actual \(df\) from the test; the extreme bounds are calculator approximations to infinity.
Because a t curve is symmetric around 0, the two areas in a two-sided test are equal. You can therefore find the area beyond \(|t|\) on the right and double it. Adding both tails directly is a useful check that you have used the correct boundaries.
Work the Alternative Hypothesis First
Before entering numbers, write down the alternative and identify the tail or tails. The sign of \(t\) tells you where the observed statistic lies, but it does not determine the p-value by itself. For example, an upper-tailed test with a negative \(t\) statistic has a large p-value: the observed statistic is on the opposite side from the evidence the alternative predicts.
Decide whether the test is lower-tailed, upper-tailed, or two-sided.
Use the observed one-sample t statistic and \(df=n-1\).
Enter bounds that include exactly the tail area or areas specified by the alternative.
State the probability assuming \(H_0\) is true, and describe results at least as extreme as the observed statistic in context.
Worked Examples
Worked Example: A Lower-Tailed P-Value
A fictional greenhouse randomly selects 3 seedlings from a group of 80 and measures the time, in days, until each first shows a visible new leaf. The sample mean is 18 days, and the sample standard deviation is \(\sqrt{3}\) days. The greenhouse wants to test whether the true mean time for seedlings in this group is less than 19 days. Find the p-value using tcdf.
State. Let \(\mu\) be the true mean time, in days, until seedlings in this group show a visible new leaf. The hypotheses are \(H_0:\mu=19\) days and \(H_a:\mu<19\) days. This is a lower-tailed test.
Plan and check conditions. The seedlings were randomly selected, supporting the Random condition. The 10% condition is met because \(0.10(80)=8\) and \(3\leq8\), so independence is reasonable for sampling without replacement. Since \(n=3<30\), the large-sample route is not met. The scenario states that the times in this group follow an approximately Normal distribution, supporting the Normal/Large Sample condition for this small sample.
Do. There are \(df=n-1=3-1=2\) degrees of freedom. The estimated standard error is
Thus, the observed statistic is
The alternative is lower-tailed, so find the area to the left of \(-1.00\):
Conclude. Assuming the true mean time is 19 days, the probability of obtaining a t statistic of \(-1.00\) or less is about \(0.2113\), or \(21.13\%\). This is the p-value for the test of whether the mean time is less than 19 days. It describes how surprising the sample result is under the null hypothesis; it is not the probability that \(H_0\) is true.
Worked Example: An Upper-Tailed P-Value
A fictional food lab randomly selects 3 sealed containers from a shipment of 150 and measures the volume of a liquid, in milliliters. The sample mean is 42 milliliters, and the sample standard deviation is \(\sqrt{3}\) milliliters. The lab tests whether the true mean volume in this shipment is greater than 40 milliliters.
State and plan. Let \(\mu\) be the true mean liquid volume, in milliliters, in containers from this shipment. The hypotheses are \(H_0:\mu=40\) milliliters and \(H_a:\mu>40\) milliliters, so this is an upper-tailed test. Random selection supports the Random condition. For the 10% condition, \(0.10(150)=15\), and \(3\leq15\), so independence is reasonable. Since \(n=3<30\), the large-sample route is not met; the lab states that container volumes in this shipment follow an approximately Normal distribution, supporting the Normal/Large Sample condition.
Do. The degrees of freedom are \(df=3-1=2\), and the estimated standard error is \(\sqrt{3}/\sqrt{3}=1\) milliliter. Therefore,
For the upper-tailed alternative, the p-value is the area to the right of \(2.00\):
As a check on this value for \(df=2\), the right-tail area at a positive \(t\) is \(\tfrac12(1-t/\sqrt{t^2+2})\). Substituting \(t=2\) gives \(\tfrac12(1-2/\sqrt{6})\approx0.0918\), matching tcdf.
Conclude. Assuming the true mean volume is 40 milliliters, the probability of obtaining a t statistic of \(2.00\) or greater is about \(0.0918\). This is the p-value for the claim that the mean volume exceeds 40 milliliters. The observed statistic is positive, in the direction of the alternative, so the p-value is the right-tail area.
Worked Example: A Two-Tailed P-Value
A fictional recreation center randomly selects 3 members from a group of 100 and records how many hours each spends using a climbing wall in a typical week. The sample mean is 6.5 hours, and the sample standard deviation is \(\sqrt{3}\) hours. The center tests whether the true mean differs from 5 hours per week.
State and plan. Let \(\mu\) be the true mean weekly climbing-wall use, in hours, for members of this group. The hypotheses are \(H_0:\mu=5\) hours and \(H_a:\mu\ne5\) hours, so this is a two-sided test. The members were randomly selected, supporting the Random condition. The 10% condition is met because \(0.10(100)=10\) and \(3\leq10\), so independence is reasonable. Since \(n=3<30\), the large-sample route is not met. The center states that weekly use in this group is approximately Normally distributed, supporting the Normal/Large Sample condition.
Do. Here \(df=3-1=2\), and the estimated standard error is \(\sqrt{3}/\sqrt{3}=1\) hour. The observed statistic is
Because the alternative allows a mean either lower or higher than 5 hours, include both tails beyond \(-1.50\) and \(1.50\). Using symmetry, double the area to the right of \(1.50\):
Conclude. Assuming the true mean weekly use is 5 hours, the probability of obtaining a t statistic at least as far from 0 as \(1.50\), in either direction, is about \(0.2724\). This is the two-sided p-value. Doubling one tail works here because the t curve is symmetric and the two-sided alternative treats departures below and above the null value equally.
Check the Direction Before You Calculate
The examples show that identical calculator syntax can produce very different areas depending on the bounds. Suppose an upper-tailed test has \(t=-1.00\) and \(df=2\). The correct p-value is the area to the right of \(-1.00\), not the area to its left:
This p-value is large because a statistic of \(-1.00\) is not in the direction of an upper-tailed alternative. Using the lower-tail command instead would give about \(0.2113\), but that would answer a different question. Always let \(H_a\), not the sign of \(t\) alone, determine the area.
A p-value is a probability under the assumption that \(H_0\) is true. It is not the probability that the null hypothesis is true, nor is it the probability that chance alone produced the data. If a significance level \(\alpha\) is specified, compare the p-value with that level to make the test decision. Without a specified \(\alpha\), report the p-value and describe the evidence without inventing a decision threshold.
Common Mistakes and AP Exam Tips
- Using the wrong tail. For \(H_a:\mu<\mu_0\), use the area left of \(t\); for \(H_a:\mu>\mu_0\), use the area right of \(t\). For \(H_a:\mu\ne\mu_0\), include both tails.
- Doubling every p-value. Double a one-tail area only for a two-sided test, and use the area beyond \(|t|\). Do not double a one-sided p-value.
- Using the wrong degrees of freedom. For a one-sample t test, use \(df=n-1\), not \(n\). The calculator needs this value because it determines which t curve supplies the area.
- Reading a calculator area without explaining it. A number alone is incomplete. Say that it is the probability, assuming \(H_0\) is true, of a statistic at least as extreme as the observed one in the direction or directions of \(H_a\).
- Calling the p-value the chance that \(H_0\) is true. It is calculated under the assumption that \(H_0\) is true; it does not measure the probability that the hypothesis itself is true.
For a clear AP response, show the alternative, \(t\), \(df\), the tcdf bounds, and the resulting area. Then state the p-value in context using a complete sentence. If you are asked to make a test decision, use the stated significance level and explain whether the result provides convincing evidence for the alternative.
Check Your Understanding
For each question, identify the appropriate tcdf area or explain what the p-value means.
- A one-sample t test has \(n=14\), \(t=-1.8\), and \(H_a:\mu<\mu_0\). What \(df\) and calculator bounds should be used?
- A test has \(t=1.4\), \(df=9\), and \(H_a:\mu>\mu_0\). Should tcdf find the area to the left or right of \(1.4\)?
- A test has \(t=-2.0\), \(df=11\), and \(H_a:\mu\ne\mu_0\). What are the two tail boundaries?
- For an upper-tailed test, the observed \(t\) statistic is negative. Why might the p-value be large?
- In context, what does a p-value of \(0.04\) mean when \(H_0\) is assumed true?