Tutorials › AP Statistics › Deciding a t Test by Comparing the P-Value to Alpha

One-sample t hypothesis tests · Tutorial 670 of 1000

Deciding a t Test by Comparing the P-Value to Alpha

Use the p-value from a t test about a population mean to make and explain a decision at the 5% significance level.

Intermediate 9 min read

What You'll Learn

  • Apply the reject-or-fail-to-reject rule at alpha = 0.05.
  • Decide using exact t-test p-values and p-value bounds from a t table.
  • See how sample means and standard deviations determine a test statistic and decision.
  • Connect a two-sided 5% t test to its matching 95% confidence interval.
  • Avoid treating failure to reject as proof that the null hypothesis is true.

What Does It Mean to Compare a P-Value with Alpha?

In “Using the t Table to Bound a P-Value,” you learned how a t table can place a p-value between two bounds. Now use the p-value to make a decision. The significance level, written \(\alpha\), is the cutoff chosen for deciding whether the sample provides convincing evidence against the null hypothesis. Here we use \(\alpha=0.05\).

The decision rule is simple: if the p-value is at most \(\alpha\), reject \(H_0\). If the p-value is greater than \(\alpha\), fail to reject \(H_0\). Rejecting means the data provide convincing evidence against the null hypothesis in favor of the alternative. Failing to reject means the data do not provide convincing evidence against the null hypothesis. It does not prove that \(H_0\) is true.

Decision rule: At significance level \(\alpha=0.05\), reject \(H_0\) when \(p\leq0.05\). When \(p>0.05\), fail to reject \(H_0\). Make the comparison using the p-value for the stated alternative hypothesis.

The p-value and alpha have different roles. The p-value is calculated from the sample and describes how unusual the observed result, or a more extreme result, would be if \(H_0\) were true. Alpha is the decision threshold chosen before interpreting the test result. A p-value is not the probability that \(H_0\) is true, and it does not tell you how large or important an effect is.

A Four-Step Decision Process

A decision is meaningful only after the test is set up correctly. As in “The One-Sample t Test Statistic” and “Finding a P-Value Using tcdf,” identify the population mean, state the hypotheses, and use the alternative hypothesis to determine which tail area is the p-value. Check the conditions for a one-sample t procedure as in “Complete Conditions Check for a Mean Inference Problem.” Then compare the resulting p-value with alpha.

1
State.
Define the population mean \(\mu\) in context and write \(H_0\) and \(H_a\). State the significance level, here \(\alpha=0.05\).
2
Plan and check conditions.
Identify the one-sample t test. Check the Random condition, the 10% condition when sampling without replacement, and the Normal/Large Sample condition.
3
Do.
Calculate \(t=(\bar{x}-\mu_0)/(s/\sqrt{n})\), use \(df=n-1\), and find the p-value for the stated alternative. Compare \(p\) with \(\alpha\).
4
Conclude.
State “reject” or “fail to reject” \(H_0\), then say whether the data provide convincing evidence for the alternative in the context of the problem.

The conditions support the use of the t procedure; they do not change the decision rule. If the procedure is not justified, a small calculator p-value alone does not repair that problem.

Worked Examples

Worked Example: A Two-Sided P-Value Below 0.05

A fictional parts facility randomly selects 16 metal pieces from a shipment of 200. Their mean mass is \(\bar{x}=52.5\) grams and their sample standard deviation is \(s=4\) grams. The facility wants to know whether the true mean mass of pieces in this shipment differs from 50 grams. Assume the population distribution of mass is approximately Normal. Use \(\alpha=0.05\).

State. Let \(\mu\) be the true mean mass, in grams, of the metal pieces in this shipment. The hypotheses are \(H_0:\mu=50\) grams and \(H_a:\mu\ne50\) grams. The significance level is \(\alpha=0.05\).

Plan and check conditions. This is a one-sample t test for a population mean. The pieces were randomly selected, supporting the Random condition. For the 10% condition, \(0.10(200)=20\), and \(16\leq20\), so independence is reasonable. Since \(n=16<30\), the large-sample route is not met; the stated approximately Normal population supports the Normal/Large Sample condition. Use \(df=16-1=15\).

Do. The estimated standard error is \(s/\sqrt{n}=4/\sqrt{16}=1\) gram. The test statistic is

$$ t=\frac{\bar{x}-\mu_0}{s/\sqrt{n}} =\frac{52.5-50}{4/\sqrt{16}} =\frac{2.5}{1} =2.50. $$

Because the alternative is two-sided, results at least as far from 0 as \(2.50\) count in both tails. Using a calculator with \(df=15\), the p-value is \(2\operatorname{tcdf}(2.5,1\mathrm{E}99,15)\approx0.02451\), rounded to five decimal places. Since \(0.02451<0.05\), reject \(H_0\).

Conclude. The sample provides convincing evidence that the true mean mass of pieces in this shipment differs from 50 grams.

This decision also agrees with the matching 95% one-sample t confidence interval. Using \(t^*=2.131\) for \(df=15\), the interval is

$$ 52.5\pm2.131\left(\frac{4}{\sqrt{16}}\right) =52.5\pm2.131 =(50.369,\ 54.631)\text{ grams}. $$

The claimed mean of 50 grams is outside this interval, just as the two-sided test at \(\alpha=0.05\) rejects \(H_0\). This agreement applies to a two-sided 5% test and its matching 95% t interval, as discussed in “Using a Confidence Interval to Test a Claimed Mean.”

Worked Example: A Lower-Tailed P-Value Below 0.05

A fictional community garden randomly selects 25 seedlings from a group of 400. The mean time for the seedlings to reach a specified growth stage is \(\bar{x}=18.4\) days, with \(s=2\) days. The gardeners want to know whether the true mean time is less than 19.2 days. Assume the population distribution is approximately Normal. Use \(\alpha=0.05\).

State. Let \(\mu\) be the true mean number of days for seedlings in this group to reach the specified growth stage. The hypotheses are \(H_0:\mu=19.2\) days and \(H_a:\mu<19.2\) days.

Plan and check conditions. Use a one-sample t test. Random selection supports the Random condition. For the 10% condition, \(0.10(400)=40\), and \(25\leq40\), supporting independence. Since \(n=25<30\), the large-sample route is not met; the stated approximately Normal population supports the Normal/Large Sample condition. The degrees of freedom are \(df=25-1=24\).

Do. The standard error is \(2/\sqrt{25}=0.4\) days. The test statistic is

$$ t=\frac{18.4-19.2}{2/\sqrt{25}} =\frac{-0.8}{0.4} =-2.00. $$

The alternative is lower-tailed, so the p-value is the area to the left of \(-2.00\). With \(df=24\), a calculator gives \(\operatorname{tcdf}(-1\mathrm{E}99,-2,24)\approx0.02847\), rounded to five decimal places. Since \(0.02847<0.05\), reject \(H_0\).

Conclude. The sample provides convincing evidence that the true mean time for these seedlings to reach the specified growth stage is less than 19.2 days.

Worked Example: A P-Value Greater Than 0.05

A fictional recreation center randomly selects 16 members from a group of 250 and records their weekly hours of exercise. The sample mean is \(\bar{x}=25\) hours and the sample standard deviation is \(s=2\) hours. The center wants to test whether the true mean weekly exercise time differs from 25 hours. Assume the population distribution is approximately Normal. Use \(\alpha=0.05\).

State. Let \(\mu\) be the true mean weekly exercise time, in hours, for members of this group. The hypotheses are \(H_0:\mu=25\) hours and \(H_a:\mu\ne25\) hours.

Plan and check conditions. Use a one-sample t test. The members were randomly selected, supporting the Random condition. For the 10% condition, \(0.10(250)=25\), and \(16\leq25\), so independence is reasonable. Since \(n=16<30\), the large-sample route is not met; the stated approximately Normal population supports the Normal/Large Sample condition. Use \(df=16-1=15\).

Do. The standard error is \(2/\sqrt{16}=0.5\) hours, so

$$ t=\frac{25-25}{2/\sqrt{16}} =\frac{0}{0.5} =0. $$

For a two-sided test, the p-value includes outcomes at least as far from 0 as the observed statistic. Since the observed statistic is 0, the p-value is \(1.0000\). Because \(1.0000>0.05\), fail to reject \(H_0\).

Conclude. The sample does not provide convincing evidence that the true mean weekly exercise time for members of this group differs from 25 hours. This result does not establish that the population mean equals 25 hours; it says that this test did not find convincing evidence of a difference.

Using a P-Value Range from a t Table

A t table may give a range rather than an exact p-value. You can still make a decision if the entire range is on one side of \(\alpha\). For example, if the table shows \(0.02<p<0.05\), the p-value is below 0.05, so reject \(H_0\). If it shows \(0.05<p<0.10\), fail to reject \(H_0\).

If a reported range straddles the cutoff—for example, \(0.04<p<0.08\)—that range alone does not determine the decision at \(\alpha=0.05\). Use a calculator or a more precise table to find out which side of 0.05 the p-value falls on. Do not choose a decision based only on the midpoint of the range.

Key idea: Compare the p-value—not the t statistic itself—with \(\alpha\). A t table gives enough information for a decision only when the p-value bounds establish whether \(p\leq0.05\) or \(p>0.05\).

Common Mistakes and AP Exam Tips

  • Reversing the decision rule. A small p-value leads to rejecting \(H_0\); a large p-value leads to failing to reject \(H_0\). State the comparison explicitly, such as “\(0.02847<0.05\), so reject \(H_0\).”
  • Writing “accept \(H_0\).” A test does not establish that the null hypothesis is true. Use “fail to reject \(H_0\)” when the p-value exceeds alpha.
  • Ignoring the alternative hypothesis. A one-sided and a two-sided test can have different p-values for the same t statistic. Use the p-value for the alternative actually stated.
  • Comparing alpha with the t statistic. The t statistic is measured in standard-error units; alpha is a probability cutoff. Compare alpha with the p-value, not with \(t\).
  • Rounding before deciding. If a p-value is close to 0.05, keep enough digits to make the comparison. A displayed value of 0.050 may hide a value slightly above or below the cutoff.
  • Confusing statistical evidence with practical importance. Rejecting \(H_0\) indicates convincing evidence against the null at the selected level; it does not by itself show that the difference is large or important.
  • Leaving out the context. A full-credit response names the population mean and states what the decision says about the alternative in the situation. Avoid a conclusion that says only “significant” or “not significant.”
Key takeaway: For a t test about a population mean at \(\alpha=0.05\), reject \(H_0\) when the correctly calculated p-value is at most 0.05; otherwise, fail to reject \(H_0\). Give the decision in context, and do not treat failure to reject as proof that the null is true.

Check Your Understanding

For each decision, use \(\alpha=0.05\). When asked for a conclusion, state it in context without claiming that \(H_0\) has been proved.

  1. A one-sample t test has \(p=0.041\). Should you reject or fail to reject \(H_0\)?
  2. A one-sample t test has \(p=0.072\). What decision should you make, and what does it not establish?
  3. A two-sided one-sample t test has \(n=16\), \(\bar{x}=31\), \(s=4\), and tests \(H_0:\mu=30\). Calculate \(t\) and \(df\). Which tail areas belong in its p-value?
  4. A t table gives \(0.03<p<0.06\). Is that range sufficient to decide at \(\alpha=0.05\)? Explain what you would do next.
  5. For a two-sided 5% one-sample t test, a 95% t confidence interval contains the claimed mean. What is the matching test decision?