One P-Value, Several Possible Decisions
In “Making a Decision From P-Value and Alpha,” you learned to compare a p-value with the significance level chosen for a test. Here we focus on what happens when that same p-value is compared with different possible values of alpha. The test result does not change, but the decision can.
A significance level of \(\alpha=0.10\) sets a less strict threshold for rejecting \(H_0\) than \(\alpha=0.05\), which in turn is less strict than \(\alpha=0.01\). For a fixed p-value, moving to a smaller alpha can change a rejection into a failure to reject. It cannot change the p-value itself.
When the p-value falls between two possible alpha values, the decision differs at those levels. For example, if \(0.01<p\le0.05\), the decision is to reject at \(\alpha=0.05\) but fail to reject at \(\alpha=0.01\). If \(p\le0.01\), it is at or below all three thresholds considered here, so the decision is to reject at 0.10, 0.05, and 0.01. These comparisons describe how a result relates to different thresholds; they do not authorize changing the planned threshold to get a preferred result.
Worked Example: The Same P-Value at All Three Levels
A fictional study compares the time needed to complete a digital scheduling task using two interfaces. Researchers take independent random samples of students from the two defined student groups. A two-sample t test for a difference in mean completion times reports a two-sided p-value of \(0.032\), rounded. The test’s conditions have been checked and are considered reasonable. Suppose we compare this one result with \(\alpha=0.10\), \(\alpha=0.05\), and \(\alpha=0.01\).
Let \(\mu_1\) and \(\mu_2\) be the true mean task-completion times, in seconds, for all students in the two defined groups using interface 1 and interface 2, respectively. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\).
Use an unpooled two-sample t test. The two groups were selected as independent random samples, supporting inference to the defined student populations. The samples are each less than 10% of their respective populations, so the 10% condition is reasonable. Each group’s data are reasonably compatible with a t procedure, with no serious skewness or outliers. The two-sided alternative matches the question about a difference in either direction.
At \(\alpha=0.10\), \(0.032\le0.10\), so reject \(H_0\). At \(\alpha=0.05\), \(0.032\le0.05\), so reject \(H_0\). At \(\alpha=0.01\), \(0.032>0.01\), so fail to reject \(H_0\).
If \(\alpha=0.10\) or \(\alpha=0.05\) had been selected in advance, the data would provide convincing evidence that the true mean completion times differ between the two defined student populations. If \(\alpha=0.01\) had been selected in advance, the data would not provide convincing evidence of a difference at that level.
The samples were randomly selected, so the conclusions above refer to the defined student populations from which they were drawn. The p-value remains \(0.032\) in every comparison; only the threshold changes. The study should use its preselected alpha for its primary decision, not select among these three conclusions after seeing the result.
| Alpha compared with the same p-value | Comparison | Decision |
|---|---|---|
| \(\alpha=0.10\) | \(0.032\le0.10\) | Reject \(H_0\) |
| \(\alpha=0.05\) | \(0.032\le0.05\) | Reject \(H_0\) |
| \(\alpha=0.01\) | \(0.032>0.01\) | Fail to reject \(H_0\) |
This table shows a useful threshold fact: a p-value of \(0.032\) would lead to rejection for any alpha at least \(0.032\), and to failure to reject for any alpha smaller than \(0.032\). The decision rule includes equality, so an exact p-value of \(0.032\) compared with \(\alpha=0.032\) would result in rejection.
How the Decision Changes Across the Thresholds
The possible decisions fall into several ranges. For the three alpha values in this tutorial, the comparisons can be summarized as follows. The endpoints matter: equality with alpha leads to rejection.
| Range containing the p-value | At \(\alpha=0.10\) | At \(\alpha=0.05\) | At \(\alpha=0.01\) |
|---|---|---|---|
| \(p\le0.01\) | Reject | Reject | Reject |
| \(0.01<p\le0.05\) | Reject | Reject | Fail to reject |
| \(0.05<p\le0.10\) | Reject | Fail to reject | Fail to reject |
| \(p>0.10\) | Fail to reject | Fail to reject | Fail to reject |
This classification is a way to understand the decision rule, not a substitute for selecting alpha before examining the data. In particular, a p-value between 0.05 and 0.10 is not a reason to promote \(\alpha=0.10\) after seeing the result. If 0.05 was planned, use 0.05.
Worked Example: A Result Between 0.05 and 0.10
A fictional public library takes a random sample of patrons and tests whether the true mean time patrons spend using a computer differs from 40 minutes. A valid two-sided one-sample t test reports \(p=0.084\), rounded. The sample is less than 10% of the target population, and its distribution is reasonably compatible with a t procedure. Compare the result with each proposed threshold.
Let \(\mu\) be the true mean computer-use time, in minutes, for patrons in the target population. The hypotheses are \(H_0:\mu=40\) minutes and \(H_a:\mu\ne40\) minutes.
At \(\alpha=0.10\), \(0.084\le0.10\), so reject \(H_0\). At \(\alpha=0.05\), \(0.084>0.05\), so fail to reject \(H_0\). At \(\alpha=0.01\), \(0.084>0.01\), so fail to reject \(H_0\).
Thus, if the library had planned \(\alpha=0.10\), the data would provide convincing evidence that the true mean computer-use time differs from 40 minutes. If it had planned \(\alpha=0.05\) or \(\alpha=0.01\), the data would not provide convincing evidence of a difference at that level. Because this is a random sample, the conclusion can refer to the target patron population. The decision still does not prove that the mean is exactly 40 minutes when we fail to reject.
Worked Example: A Result Below All Three Thresholds
A fictional recreation center randomly samples members to test whether the true mean time members spend exercising during a visit exceeds 50 minutes. A valid one-sided t test reports \(p=0.006\), rounded. The sample is less than 10% of the membership, and the sample data show no serious skewness or outliers.
Let \(\mu\) be the true mean exercise time, in minutes, per visit for members of this recreation center. The hypotheses are \(H_0:\mu=50\) minutes and \(H_a:\mu>50\) minutes. The alternative is one-sided because the research question asks whether the mean exceeds 50 minutes.
At \(\alpha=0.10\), \(0.006\le0.10\), so reject \(H_0\). At \(\alpha=0.05\), \(0.006\le0.05\), so reject \(H_0\). At \(\alpha=0.01\), \(0.006\le0.01\), so reject \(H_0\). At each of these levels, the data provide convincing evidence that the true mean exercise time per visit for this center’s members exceeds 50 minutes.
Here, the conclusion is the same at all three thresholds because the p-value is smaller than each one. This does not mean the evidence or the p-value has been recalculated three times. It is one test result compared with three different decision rules.
Why Alpha Must Be Chosen in Advance
Alpha is chosen before examining the data because the threshold affects the chance of rejecting a true null hypothesis over repeated uses of a test. A larger alpha, such as 0.10 rather than 0.01, allows rejection for a wider range of p-values. That more permissive rule also allows more opportunities for a Type I error: rejecting \(H_0\) when it is true. This is a long-run property of the decision procedure, not the probability that a particular rejection is wrong.
Reporting how a p-value compares with several alpha values can help readers see how sensitive the reject-or-fail-to-reject decision is to the threshold. But the report should distinguish that comparison from the primary conclusion. For example: “The planned significance level was 0.05; because \(p=0.032\le0.05\), we reject \(H_0\). At 0.01, we would fail to reject.” Do not write as if the study had planned all three levels if it had only specified one.
A change in decision does not turn a p-value into a probability that the null hypothesis is true. Nor does rejection by itself establish that a mean difference is large or important in practice. A decision describes whether the test met the preselected evidence threshold; interpretation should still identify the population mean or means and the claim being tested.
Common Mistakes and AP Exam Tips
- Changing alpha to match the result: Do not choose 0.10 only because a p-value exceeds 0.05, or switch to 0.01 to avoid rejecting. A full-credit response uses the alpha selected before examining the data.
- Changing the p-value between comparisons: Keep the reported p-value fixed. Compare it separately with 0.10, 0.05, and 0.01; do not treat each threshold as a new test calculation.
- Reversing the inequality: Reject when \(p\le\alpha\), including equality. If \(p>\alpha\), fail to reject. Writing the comparison before the decision helps prevent this error.
- Writing only “reject” or “fail to reject”: Name the claim in the alternative hypothesis and state what the decision means in context. For a two-sided test, describe evidence of a difference; for a one-sided test, describe evidence in the specified direction.
- Saying “accept the null” or “prove no difference”: Failure to reject means the data did not provide convincing evidence for the alternative at the chosen level. It does not establish that the null value is true.
- Generalizing beyond the study design: Random assignment supports a causal comparison for the units in an experiment, but it does not by itself make those units representative of a wider population. Generalization requires appropriate random sampling from that population.
- Ignoring rounding near alpha: If the displayed p-value is very close to a threshold, use additional available digits. A rounded value may hide whether the unrounded p-value is just above or below alpha.
A concise, complete decision sentence includes the comparison, the preselected alpha, and a contextual conclusion. For example: “Because the p-value of \(0.032\) is less than the preselected \(\alpha=0.05\), reject \(H_0\). The data provide convincing evidence that the true mean completion times differ between the two defined student populations.” If the p-value is greater than alpha, say “fail to reject” and that the data do not provide convincing evidence for the alternative claim at that level.
Check Your Understanding
For each situation, compare the same p-value with the stated significance levels and write an appropriate conclusion when the context is provided.
- A valid two-sided test about two population means has \(p=0.032\). What is the decision at \(\alpha=0.10\), \(0.05\), and \(0.01\)?
- A random sample of residents is used in a valid test of whether the true mean weekly recycling mass exceeds a target. The test reports \(p=0.084\). What are the decisions at the three significance levels, and at which level would the data provide convincing evidence for the alternative?
- A valid test has \(p=0.010\) exactly. State the decision at \(\alpha=0.01\), and explain why equality matters.
- A study planned \(\alpha=0.05\) and obtained \(p=0.061\). Why is it inappropriate to change the planned alpha to 0.10 after seeing the p-value?
- An experiment randomly assigns its participants to two methods and finds convincing evidence of a difference in mean response. What does random assignment support, and what does it not establish about generalizing to all people in a wider population?