A Test Decision Is Not the Same as Knowing the Truth
A significance test gives a decision about a null hypothesis, but the test does not tell us whether that hypothesis is actually true. This is why a decision can be correct or can lead to an error. In “Two Possible Errors in a Significance Test,” we named the two errors: rejecting a true null hypothesis is a Type I error, and failing to reject a false null hypothesis is a Type II error. Here, we connect those definitions to the wording of a test decision in context.
The key is to consider two things separately: the decision made from the sample and the actual truth about the population. The sample determines whether we reject or fail to reject \(H_0\). The actual population proportion determines whether that decision was correct. In a real investigation, that truth is generally unknown; we describe the error that could have occurred, not one that we know did occur.
For a test about a population proportion \(p\), read the decision in context before naming the possible error. If the alternative is \(H_a:p>p_0\), rejecting the null means the sample provides evidence that the population proportion is greater than \(p_0\). A possible Type I error would be concluding that it is greater when the null model is true. Failing to reject means the sample did not provide convincing evidence for that increase; a possible Type II error would be failing to find evidence of an increase when the true proportion is greater than \(p_0\).
The same logic applies to a two-sided alternative or a claim that the proportion is lower. The direction of the claim changes the contextual wording, but not the decision-error pairing. Use the decision-by-truth table as a quick check.
| Test decision | If \(H_0\) is true | If \(H_0\) is false |
|---|---|---|
| Reject \(H_0\) | Type I error | Correct decision |
| Fail to reject \(H_0\) | Correct decision | Type II error |
This table does not tell us which outcome happened in a particular study. It tells us which outcome is possible under each combination of decision and truth. When writing a conclusion, avoid wording that treats a possible error as a known fact.
A Contextual Error-Identification Routine
After deciding a test, use a short routine to describe the relevant possible error accurately. This builds on the four-step testing approach from earlier in the course: the State and Plan identify the question and procedure, the Do gives the statistical result, and the Conclude step interprets the decision in context.
Identify what \(H_0\) and \(H_a\) say about the population parameter in the setting.
Use “reject \(H_0\)” or “fail to reject \(H_0\),” based on the p-value and the chosen significance level.
For a rejection, ask what Type I error would mean if the null were true. For a failure to reject, ask what Type II error would mean if the alternative were true.
Say that the result “could be” or “would be” an error under the specified truth. Do not claim that an error actually occurred unless the population truth is known.
For a complete explanation, name the population and outcome, describe the incorrect conclusion or missed effect, and state the condition under which it would be an error. For example, “A Type I error would be concluding that the proportion of residents who support the proposal is greater than 0.40 when the true proportion is 0.40.” The condition is essential: a rejection is not automatically a Type I error.
Worked Example: Rejecting a Null Hypothesis About Bike-Helmet Use
A health educator wants to know whether more than 55% of teenagers in a district regularly wear a bike helmet. A random sample of 120 teenagers is selected from the district’s 4,000 teenagers, and 78 report regular helmet use. At \(\alpha=0.05\), test \(H_0:p=0.55\) against \(H_a:p>0.55\), where \(p\) is the proportion of all teenagers in the district who regularly wear a helmet. Then identify the possible error connected to the decision.
State: The question is whether the district’s proportion of teenagers who regularly wear a helmet is greater than 0.55. The null hypothesis is \(H_0:p=0.55\), and the alternative is \(H_a:p>0.55\).
Plan and conditions: Use a one-proportion \(z\) test. The sample is stated to be random. The 10% condition holds because \(120\le 0.10(4{,}000)=400\), supporting independence when sampling without replacement. Under the null model, the expected number of successes is \(np_0=120(0.55)=66\), and the expected number of failures is \(n(1-p_0)=120(0.45)=54\). Both are at least 10, so the Large Counts condition is met.
Do: The sample proportion is \(\hat{p}=78/120=0.65\). For the test, calculate the standard deviation of the sample proportion under \(H_0\) and then the \(z\) statistic:
Because the alternative is greater than, the p-value is the upper-tail probability: \(P(Z\ge 2.202)\approx 0.0138\), rounded. Since \(0.0138<0.05\), reject \(H_0\).
Conclude: The sample provides convincing evidence that more than 55% of teenagers in this district regularly wear a bike helmet. A possible Type I error would be concluding that the proportion is greater than 55% when the true proportion is 55% under the null model. The test result does not establish that such an error occurred; if the true proportion is greater than 55%, the rejection is a correct decision.
Worked Example: Failing to Reject a Null Hypothesis About a Library Service
A library is evaluating whether more than half of its customers use its online book-reservation service. From a random sample of 100 customers, 56 say they use the service. The library has 4,000 customers. At \(\alpha=0.05\), test \(H_0:p=0.50\) against \(H_a:p>0.50\), where \(p\) is the proportion of all customers who use the service. Identify the possible error connected to the decision.
State: The question is whether more than half of the library’s customers use the online service. We test \(H_0:p=0.50\) against \(H_a:p>0.50\).
Plan and conditions: Use a one-proportion \(z\) test. The sample is random, and the 10% condition holds because \(100\le 0.10(4{,}000)=400\). Under the null model, \(np_0=100(0.50)=50\) customers are expected to use the service and \(n(1-p_0)=100(0.50)=50\) are expected not to use it. Both expected counts meet the Large Counts condition.
Do: The sample proportion is \(\hat{p}=56/100=0.56\). The standard deviation of the sample proportion under the null model and the test statistic are
The upper-tail p-value is \(P(Z\ge1.20)\approx0.1151\), rounded. Since \(0.1151>0.05\), fail to reject \(H_0\).
Conclude: The sample does not provide convincing evidence that more than half of the library’s customers use the online reservation service. A possible Type II error would be failing to find convincing evidence that more than half use it when the true proportion is actually greater than 0.50. If the true proportion is 0.50, failing to reject is not a Type II error; it is the correct decision for that truth. The result does not prove that exactly half of the customers use the service.
Notice How the Alternative Shapes the Error in Context
A decision error must be described using the claim in the alternative hypothesis. If the alternative says “less than,” a Type I error after rejecting means concluding that the proportion is below the benchmark when the null model is true. A Type II error after failing to reject means missing a real decrease. The words “increase,” “decrease,” or “different” should match the stated alternative rather than the direction of the sample result alone.
Worked Example: Rejecting a Null Hypothesis About Equipment Repairs
A recreation center claims that no more than 20% of its rental bicycles need a repair after a day of use. A maintenance supervisor takes a random sample of 150 bicycles from the center’s 5,000 bicycles; 20 in the sample need a repair. At \(\alpha=0.05\), test \(H_0:p=0.20\) against \(H_a:p<0.20\), where \(p\) is the proportion of all the center’s rental bicycles that need a repair after a day. State the possible error connected to the decision.
State: The question is whether the proportion of rental bicycles needing a repair after a day is below 0.20. We test \(H_0:p=0.20\) against \(H_a:p<0.20\).
Plan and conditions: Use a one-proportion \(z\) test. The sample is random. The 10% condition holds because \(150\le0.10(5{,}000)=500\). Under the null model, the expected number needing repair is \(150(0.20)=30\), and the expected number not needing repair is \(150(0.80)=120\). Both are at least 10, so the Large Counts condition is met.
Do: The sample proportion needing a repair is \(\hat{p}=20/150\approx0.1333\). The standard deviation under the null model and the test statistic are
Because the alternative is less than, the p-value is the lower-tail probability: \(P(Z\le-2.041)\approx0.0206\), rounded. Since \(0.0206<0.05\), reject \(H_0\).
Conclude: The sample provides convincing evidence that less than 20% of the center’s rental bicycles need a repair after a day of use. A possible Type I error would be concluding that the proportion is below 20% when the true proportion is 20% under the null model. If the true proportion is below 20%, the rejection is a correct decision. Had the test instead failed to reject, a possible Type II error would have been missing a real decrease in the proportion.
Common Mistakes and AP Exam Tip
- Calling every rejection a Type I error: A rejection is a Type I error only if the null hypothesis is actually true. Say “a possible Type I error would be…” and state the truth under which it occurs.
- Calling every failure to reject a Type II error: A failure to reject is a Type II error only when the null is false. If the null is true, the decision is correct.
- Saying “accept the null” or “prove the null”: Use “fail to reject \(H_0\).” That wording reports the test decision without claiming that the null has been established as true.
- Describing the error without the setting: “A Type II error is failing to reject a false null” gives the general definition, but a full-credit contextual response explains what effect was missed and for which population or outcome.
- Reversing the conclusion and the error: For a test of \(H_a:p>p_0\), a Type I error is falsely concluding that \(p>p_0\); a Type II error is missing an actual increase. Check the direction in \(H_a\) before writing either statement.
- Claiming to know an error occurred from the p-value: The p-value helps determine the test decision. It does not tell us the actual population proportion, so it cannot tell us whether that decision was an error.
On an AP response, write the conclusion first in context, then describe the relevant possible error conditionally. For a rejection, connect the claim supported by the alternative to a Type I error if the null is true. For a failure to reject, describe the effect that could have been missed if the alternative is true. This shows that you understand both the statistical decision and the uncertainty about the population truth.
Check Your Understanding
For each item, distinguish the test decision from the actual truth about the population.
- A test of \(H_0:p=0.40\) against \(H_a:p>0.40\) rejects \(H_0\). Describe the possible Type I error in context if \(p\) is the proportion of residents who recycle.
- A test of \(H_0:p=0.60\) against \(H_a:p<0.60\) fails to reject \(H_0\). What actual population condition would make this a Type II error?
- Why is it incorrect to say that a failure to reject proves the null hypothesis is true?
- A test rejects \(H_0\), and the null hypothesis is in fact false. Is this a Type I error, a Type II error, or a correct decision? Explain.
- In a test with \(H_a:p\ne0.25\), what would a Type I error mean if \(p\) is the proportion of users who complete an online form?