A Test Decision Is Not the Same as Knowing the Truth
In “Making a Decision in a Chi-Square Test,” you compared a p-value with the significance level and decided either to reject \(H_0\) or to fail to reject \(H_0\). That decision is based on sample data. But the hypotheses describe a population truth, which the test does not reveal with certainty. The decision can agree with the truth, or it can be wrong.
There are two possible states of the world for a significance test: \(H_0\) is true, or \(H_0\) is false. There are also two possible decisions: reject \(H_0\), or fail to reject \(H_0\). Pairing each possible decision with each possible state gives four outcomes. Two are correct decisions, and two are errors.
The decision table keeps the two questions separate: What did the test decide? And what is actually true in the population? The second question is generally unknown in practice—that is why we conduct the test.
| Test decision | \(H_0\) is actually true | \(H_0\) is actually false |
|---|---|---|
| Reject \(H_0\) | Type I error | Correct decision |
| Fail to reject \(H_0\) | Correct decision | Type II error |
Reading across the table helps prevent a common mix-up. A Type I error is in the “reject” row and “\(H_0\) is true” column. A Type II error is in the “fail to reject” row and “\(H_0\) is false” column. The error names are defined by the combination of decision and truth, not by whether the p-value is “large” or “small” on its own.
Translate the Table Into the Situation
To use the table in context, first write what \(H_0\) says in words. Then consider each possible decision under each possible truth. For a Type I error, imagine that the null claim really is true, but the test rejects it. For a Type II error, imagine that the null claim is false, but the test does not reject it.
The phrase fail to reject matters. If the p-value is greater than \(\alpha\), the data have not provided convincing evidence against \(H_0\). That does not establish that \(H_0\) is true. If \(H_0\) is in fact false, the test may simply have failed to detect the difference or association; that possible outcome is a Type II error.
The opposite is also important. Rejecting \(H_0\) means the data provide convincing evidence against the null model according to the test and chosen significance level. It does not guarantee that \(H_0\) was false. If the null is actually true, rejecting it is a Type I error.
Worked Examples: Reading the Four Outcomes
Worked Example: A New Reminder for Clinic Appointments
A clinic tests whether a new text-message reminder changes the proportion of patients who miss their appointments. Let \(p\) be the proportion of all eligible patients who would miss an appointment with the new reminder, and let \(p_0\) be the corresponding proportion under the clinic’s existing process. The hypotheses are \(H_0:p=p_0\) and \(H_a:p\ne p_0\).
Suppose the test rejects \(H_0\). The decision is now known, but the actual population truth is not. If the new reminder really does not change the missed-appointment proportion, then \(H_0\) is true and the rejection is a Type I error. If the reminder really does change the proportion, then \(H_0\) is false and rejecting it is a correct decision.
Now suppose instead that the test fails to reject \(H_0\). If the reminder truly does not change the proportion, the decision is correct. If it truly does change the proportion, the test has failed to detect that change; this is a Type II error.
The table describes all four possibilities without claiming that the clinic knows which one occurred. In particular, a non-significant result cannot be reported as proof that the reminder has no effect.
Worked Example: Checking a Bottle-Filling Station
A school facilities team investigates whether a bottle-filling station dispenses the target volume of water on average. Let \(\mu\) be the station’s true mean volume per fill, in milliliters. The hypotheses are \(H_0:\mu=500\) milliliters and \(H_a:\mu\ne500\) milliliters.
Imagine first that the station’s true mean is exactly 500 milliliters. In that case, \(H_0\) is true. If the sample test nevertheless rejects \(H_0\), it concludes that the evidence indicates a difference when there is no population difference: a Type I error. If the test fails to reject \(H_0\), that is a correct decision.
Now imagine the station’s true mean is 492 milliliters. Then \(H_0\) is false. If the test rejects \(H_0\), that is a correct decision because the test detects evidence that the true mean differs from 500 milliliters. If the test fails to reject \(H_0\), the decision misses the real difference: a Type II error.
Notice that “correct decision” does not require the test to identify the exact true mean. In this example, the test is evaluating whether the mean equals 500 milliliters or differs from it. The table classifies the decision relative to that question.
Worked Example: Comparing Two Garden Treatments
A community garden compares two treatments by testing whether the distribution of seedling condition (strong, average, or weak) is the same in both treatment groups. Let \(H_0\) state that the condition distributions are the same, and let \(H_a\) state that they differ. This is a chi-square test of homogeneity, as in “Chi-Square Test of Homogeneity Full Worked Example.”
Suppose the test’s p-value is greater than the chosen \(\alpha\), so the team fails to reject \(H_0\). If the distributions really are the same, this is a correct decision. If they actually differ, the test has not detected that difference, so this outcome is a Type II error.
For the other decision, suppose the p-value is at most \(\alpha\), so the team rejects \(H_0\). If the distributions really differ, rejecting is a correct decision. If the distributions are actually the same, rejecting is a Type I error.
The test conclusion should still be stated in context: either the data provide convincing evidence that the seedling-condition distributions differ between treatments, or they do not provide convincing evidence of a difference. The decision table adds a separate point: because the actual population distributions are unknown, the team cannot tell from the test alone whether its particular decision is correct.
What the Significance Level Tells You
Before collecting data, researchers choose a significance level, \(\alpha\), such as 0.05. The decision rule is set up so that, when the test’s assumptions and procedure are appropriate, the probability of a Type I error is controlled at the chosen level. In repeated use of the procedure when \(H_0\) is true, the long-run rate of Type I errors is at most \(\alpha\).
This does not mean that after a particular test rejects \(H_0\), there is a 5% chance that the rejection is an error when \(\alpha=0.05\). The significance level describes the procedure’s long-run error behavior under the assumption that \(H_0\) is true; it is not the probability that \(H_0\) is true or false after seeing the data.
A Type II error has a different structure. Its probability depends on which alternative is actually true, as well as factors such as sample size and how far the true population value is from the null value. For that reason, one test does not generally have a single Type II error probability that applies equally to every possible false null. Later tutorials will examine these error ideas further, including how they are defined in context.
Common Mistakes and AP Exam Tip
- Swapping the two errors: Type I means reject a true \(H_0\). Type II means fail to reject a false \(H_0\). Locate both the decision and the actual truth in the table before naming an error.
- Treating “fail to reject” as “accept”: A large p-value does not prove the null hypothesis. A careful response says the data do not provide convincing evidence against \(H_0\), not that \(H_0\) has been shown true.
- Assuming a significant result cannot be wrong: A rejection can be a Type I error if \(H_0\) is actually true. Statistical significance is a decision under a procedure, not certainty about the population.
- Calling every wrong decision a Type I error: The error type depends on which truth-decision combination occurred. Rejecting a false null is correct; failing to reject a false null is Type II.
- Claiming the test tells us which outcome occurred: We know whether the test rejected or failed to reject. We usually do not know the actual population truth, so we cannot identify the realized outcome with certainty.
- Misinterpreting \(\alpha\): With \(\alpha=0.05\), do not say there is a 5% probability that the null is true or a 5% probability that this particular rejection is wrong. Relate \(\alpha\) to the long-run probability of a Type I error when \(H_0\) is true.
For a full-credit explanation in context, describe the null claim, name the test decision, and state what actual population condition would make that decision an error. For example: “A Type I error would occur if the reminder truly does not change the missed-appointment proportion, but the test rejects the null claim of no change.”
Check Your Understanding
For each item, use the decision table and explain your answer in context where requested.
- A test rejects \(H_0\), and \(H_0\) is actually true. Name the outcome.
- A test fails to reject \(H_0\), and \(H_0\) is actually false. Name the outcome.
- A test fails to reject a null hypothesis that is in fact true. Is this a correct decision or an error?
- A test of a new irrigation method rejects the null hypothesis of no change in the proportion of plants that survive. Describe the condition that would make this a Type I error.
- Explain why a result that fails to reject \(H_0\) does not prove \(H_0\) is true.
- In one sentence, explain what \(\alpha=0.05\) says about the testing procedure and what it does not say about a particular result.