From a P-Value to a Test Decision
In Interpreting a P-Value in Context, we learned to describe what a p-value says about sample results under the null hypothesis. The next step is to compare that p-value with a significance level, written \(\alpha\), and decide whether the evidence is strong enough to reject \(H_0\).
The significance level is the cutoff for deciding whether a p-value is small enough to count as convincing evidence against the null hypothesis. It is chosen before analyzing the data. Common choices include 0.10, 0.05, and 0.01. A smaller \(\alpha\) sets a stricter cutoff: the p-value must be smaller to lead to rejection.
This matches the convention established in Statistical Significance Versus Practical Significance: a result is statistically significant at level \(\alpha\) when the p-value is less than or equal to \(\alpha\). Most calculator results are rounded, so an exact tie is unusual in practice. If a problem gives an exact tie, use the stated convention: reject when the p-value equals \(\alpha\).
A decision is not the same as a statement about whether the null hypothesis is true. Rejecting \(H_0\) means the data provide convincing evidence in favor of the alternative hypothesis. Failing to reject \(H_0\) means the data do not provide convincing evidence against \(H_0\). It does not mean that we have proved \(H_0\) true.
How to Compare the Two Numbers
The comparison itself is simple, but it must use the p-value for the test that was actually conducted and the significance level chosen for that test. The alternative hypothesis determines the p-value’s tail or tails, as discussed in Choosing One-Sided or Two-Sided Alternatives. Once that p-value has been found, compare its numerical value with \(\alpha\).
For example, if the p-value is 0.032 and \(\alpha=0.05\), then \(0.032<0.05\), so reject \(H_0\). If the same p-value is compared with \(\alpha=0.01\), then \(0.032>0.01\), so fail to reject \(H_0\). The p-value has not changed; the decision changes because the cutoff is different.
Worked Examples: Making the Decision
The examples below focus on applying the comparison and stating its meaning. The first includes the full four-step structure of a one-proportion \(z\)-test, so the decision remains connected to the question, the conditions, and the evidence.
Worked Example: Is the Proportion Above the Benchmark?
A fictional community garden group claims that 30% of residents who receive its monthly newsletter sign up for at least one volunteer shift. A random sample of 400 newsletter recipients includes 136 who signed up. The question is whether the true proportion is greater than 0.30. Use \(\alpha=0.05\).
State: Let \(p\) be the true proportion of recipients of this group’s monthly newsletter who sign up for at least one volunteer shift. The hypotheses are \(H_0:p=0.30\) and \(H_a:p>0.30\).
Plan: Use a one-proportion \(z\)-test. The recipients were randomly sampled, so the Random condition is met for recipients represented by the sampling process. The sample was selected without replacement from 8,000 recipients, and \(400\leq0.10(8{,}000)=800\), so the 10% condition is met. Under \(H_0\), the expected number who sign up is \(np_0=400(0.30)=120\), and the expected number who do not is \(n(1-p_0)=400(0.70)=280\). Both expected counts are at least 10, so the Large Counts condition is met.
Do: The observed sample proportion, null standard error, and test statistic are:
The alternative is right-tailed. The p-value is the area to the right of \(z=1.746\), which is approximately 0.0404, rounded to four decimal places. Compare it with \(\alpha=0.05\):
Conclude: Reject \(H_0\). At the 0.05 significance level, the sample provides convincing evidence that more than 30% of recipients of this community garden group’s monthly newsletter sign up for at least one volunteer shift.
Worked Example: One P-Value, Three Significance Levels
A one-proportion test about whether a larger share of local households compost than a stated benchmark produces a p-value of 0.032. Compare this result with three possible significance levels. Assume the study’s conditions have been checked.
At \(\alpha=0.01\), \(0.032>0.01\). The decision is to fail to reject \(H_0\). The evidence is not convincing at the 0.01 level.
At \(\alpha=0.05\), \(0.032<0.05\). The decision is to reject \(H_0\). At the 0.05 level, the data provide convincing evidence in the direction of the alternative.
At \(\alpha=0.10\), \(0.032<0.10\). The decision is also to reject \(H_0\), now using the 0.10 level.
| Significance level | Comparison | Decision |
|---|---|---|
| \(\alpha=0.01\) | \(0.032>0.01\) | Fail to reject \(H_0\) |
| \(\alpha=0.05\) | \(0.032<0.05\) | Reject \(H_0\) |
| \(\alpha=0.10\) | \(0.032<0.10\) | Reject \(H_0\) |
The p-value does not change across these comparisons because the data and hypotheses are the same. The decision changes because each significance level sets a different cutoff. In an actual analysis, the significance level should be chosen before examining the data, not selected afterward to obtain a preferred conclusion.
Worked Example: A P-Value Equal to Alpha
A test of whether the true proportion of customers using a self-checkout option differs from 0.50 reports a p-value of 0.050. The test’s significance level is \(\alpha=0.050\). What is the decision?
Compare the two values:
Equality is included in the rejection rule, so reject \(H_0\). In context, at the 0.050 significance level, the data provide convincing evidence that the true proportion of customers using self-checkout differs from 0.50.
This example also shows why it is useful to state the comparison instead of relying on a vague phrase such as “the p-value is about the same as alpha.” With the stated rule, equality has a clear decision.
Worked Example: A Large P-Value
A random sample is used to test whether a greater proportion of students at a fictional high school bring a refillable water bottle than the school’s benchmark of 0.60. The test has a p-value of 0.184 and uses \(\alpha=0.05\). Assume the conditions for the one-proportion \(z\)-test have been met.
Compare the p-value with the significance level:
Therefore, fail to reject \(H_0\). The data do not provide convincing evidence, at the 0.05 significance level, that more than 60% of students at this school bring a refillable water bottle.
This decision does not establish that the true proportion is exactly 0.60. The sample may not provide enough evidence to distinguish the true proportion from the null value. A different sample could produce a different p-value.
What the Decision Does—and Does Not—Say
A decision made by comparing the p-value with \(\alpha\) answers a particular question: is the evidence against \(H_0\) strong enough under the selected cutoff? The conclusion should refer to the population parameter and the direction in \(H_a\), not treat the sample proportion as if it were the population value.
- Reject \(H_0\): Say that the data provide convincing evidence for the alternative claim, in context. Do not say that the alternative has been proven.
- Fail to reject \(H_0\): Say that the data do not provide convincing evidence for the alternative claim, in context. Do not say that the null hypothesis has been accepted or proven true.
- Keep the alternative’s direction: If \(H_a:p>p_0\), the conclusion concerns a proportion greater than \(p_0\). If \(H_a:p<p_0\), it concerns a proportion less than \(p_0\). If \(H_a:p\ne p_0\), it concerns a difference in either direction.
The decision rule does not measure how large or important a difference is. As explained in Statistical Significance Versus Practical Significance, a result can be statistically significant without being practically important. The p-value and \(\alpha\) determine the test decision; context and the size of the observed difference matter when discussing practical importance.
Common Mistakes and AP Exam Tips
- Reversing the comparison: A small p-value supports rejection, not a large one. Write the two numbers side by side and check which is smaller.
- Using “accept \(H_0\)” after a large p-value: The AP-standard decision is “fail to reject \(H_0\).” A large p-value is not proof that the null value is correct.
- Forgetting the context in the conclusion: “Reject \(H_0\)” states the decision but not what it means. Add a sentence about convincing evidence for the alternative claim and name the population and characteristic.
- Changing \(\alpha\) after seeing the result: Doing so makes the cutoff depend on the observed data. Use the significance level specified for the test.
- Treating significance as importance: A decision at a chosen cutoff does not tell you whether the difference matters in practice. Do not call a result important based only on the comparison.
- Ignoring equality: Under the standard rule, a p-value equal to \(\alpha\) leads to rejection. Do not leave a tie without a decision.
Key Takeaway
Compare the p-value with the significance level selected for the test. A p-value less than or equal to \(\alpha\) leads to rejecting \(H_0\); a p-value greater than \(\alpha\) leads to failing to reject \(H_0\). State the decision carefully in context, without claiming that either hypothesis has been proved.
Check Your Understanding
For each item, compare the p-value with \(\alpha\), make the decision, and avoid claiming that a hypothesis has been proved.
- A test has p-value 0.024 and \(\alpha=0.05\). What is the decision?
- A test has p-value 0.071 and \(\alpha=0.05\). What is the decision, and why is “accept \(H_0\)” not the preferred wording?
- A test has p-value 0.010 and \(\alpha=0.010\). What is the decision under the standard rule?
- A test’s p-value is 0.032. Explain why its decision differs at \(\alpha=0.01\) and \(\alpha=0.05\), even though the data have not changed.
- A test of whether a population proportion exceeds a benchmark leads to failing to reject \(H_0\). Write one appropriate concluding sentence in context, without treating the null hypothesis as proved.