Statistical Errors Have Real-World Consequences
In “Describing Errors in Context,” you learned how to translate a Type I or Type II error into a sentence about a population mean. This tutorial adds a practical question: if a decision is wrong, what could happen next, and who would bear the consequences? The answer can differ sharply between a medical decision, a production line, and a legal or regulatory setting.
The test itself does not measure the cost of an error. It uses sample data to assess evidence about a population mean. People making a policy or operational decision must also consider what actions follow from the test and what harms those actions—or inaction—could cause. The same kind of statistical error can therefore have very different practical importance in different settings.
A useful cost comparison asks more than “Which error sounds worse?” Consider the severity of the possible harm, how many people or units could be affected, how long the harm could continue, whether it can be reversed, and who bears the cost. A costly delay, for example, may affect patients differently from a false alarm that triggers extra testing.
These questions do not change the parameter, the hypotheses, the t procedure, or the meaning of a p-value. Nor does a small p-value tell us how harmful a decision would be. The statistical analysis and the practical judgment are related parts of decision-making, but they answer different questions.
A Cost Audit for Mean Decisions
Start with the hypotheses and the action attached to each possible test decision. Then connect each error to the real-world result. The table below is a planning tool: its entries depend on the setting, and the consequences may affect different people.
| Test decision | If \(H_0\) is true | If \(H_0\) is false |
|---|---|---|
| Reject \(H_0\) | Type I error: consider the cost of taking action when the null claim is true. | Correct rejection: consider the benefit of responding to a real difference or change. |
| Fail to reject \(H_0\) | Correct non-rejection: consider the cost of not taking action when the null claim is true. | Type II error: consider the cost of missing a real difference or change. |
An important detail is that “cost” can include more than money. It may include health risks, avoidable exposure, wasted materials, interruption of a service, lost time, or unfair treatment. The magnitude and distribution of those consequences belong to the particular decision, not to the t statistic.
Confidence intervals can help make the uncertainty about a mean more concrete. As discussed in “Using Confidence Intervals to Judge Practical Importance,” compare plausible values in the interval with meaningful values or thresholds in the original units. An interval does not assign a cost to each value, and it does not make a wrong decision impossible.
Worked Examples: Comparing the Consequences
Worked Example: A Medical Decision About Recovery Time
A fictional hospital evaluates a new care plan. Let \(\mu\) be the true mean recovery time, in days, for patients represented by the study. The hospital wants evidence that the plan reduces mean recovery time below an 8-day benchmark. A random sample of 16 patients using the plan has \(\bar{x}=7.4\) days and \(s=2.4\) days. Use a one-sample t test of \(H_0:\mu=8\) against \(H_a:\mu<8\), with \(\alpha=0.05\). Then compare the costs of the two error types.
State. The parameter is the population mean recovery time for patients using the plan. A Type I error would be concluding that the mean is below 8 days when the null claim is true. A Type II error would be failing to detect a mean below 8 days when the null claim is false.
Plan and conditions. A one-sample t test fits because one quantitative sample is being compared with a fixed benchmark. Assume the patients were selected using an appropriate random process. The observations should be independent; if sampling without replacement, 16 patients should be less than 10% of the population represented. Because \(n=16\) is small, the distribution of recovery times should have no strong skewness or outliers. The problem provides no plot, so this last condition would need to be checked with the data before relying on the test.
Do. The standard error is \(s/\sqrt{n}\), and the test statistic is the difference between the sample mean and null mean divided by that standard error:
The lower-tail p-value is approximately 0.1667, rounded. Since \(0.1667>0.05\), fail to reject \(H_0\). The data do not provide convincing evidence that the population mean recovery time with the new plan is below 8 days.
A conventional 95% confidence interval gives another view of the uncertainty. For 15 degrees of freedom, \(t^*=2.131\), so the margin of error is \(2.131(0.6)=1.2786\) days. The interval is \(7.4\pm1.2786\), or approximately \((6.121,8.679)\) days, rounded. The benchmark of 8 days is among the plausible values in this interval.
Conclude and weigh the costs. This sample does not establish that the mean is 8 days; as explained in “Why a t Test Never Proves the Null Mean,” failing to reject does not prove the null. If the true mean were 8 days, concluding that the plan reduces mean recovery time would be a Type I error. That might lead the hospital to adopt a plan that does not produce the claimed average reduction, potentially diverting resources from other care. If the true mean were, for example, 7 days, failing to reject could be a Type II error: a real average improvement would go undetected, potentially delaying a beneficial plan. Which consequence matters more depends on the size and likelihood of the harms and the costs of further evaluation—not just on this p-value.
Worked Example: Stopping a Production Line for Possible Underfilling
A fictional factory checks the fill volume of containers from a production line. Let \(\mu\) be the true mean fill, in milliliters. The target is 500 milliliters, and managers want to know whether the process is underfilling. In a random sample of 25 containers, \(\bar{x}=498.8\) milliliters and \(s=3.0\) milliliters. Test \(H_0:\mu=500\) against \(H_a:\mu<500\) at \(\alpha=0.05\).
State and plan. The parameter is the population mean fill volume from the process. Use a one-sample t test because one quantitative sample is compared with a fixed target. Assume the containers were randomly sampled and the measurements are independent. If sampling without replacement, 25 containers should be less than 10% of the production represented. With \(n=25\), check that the fill-volume distribution has no strong skewness or outliers; this should be assessed from the sample data.
Do. The standard error and test statistic are:
The lower-tail p-value is approximately 0.0285, rounded. Since \(0.0285<0.05\), reject \(H_0\). The data provide convincing evidence that the population mean fill is below 500 milliliters.
Conclude and weigh the costs. If the true mean were exactly 500 milliliters, rejecting the null and concluding that the process underfills would be a Type I error. A resulting line stoppage could waste materials, interrupt production, and require unnecessary adjustment. If the true mean were below 500 milliliters but the test failed to reject, that would be a Type II error. Underfilled containers could continue reaching customers, and the process might remain out of compliance. The two costs may fall on different groups: the company may bear downtime costs, while customers may bear the consequences of short fills. A significant result gives evidence against the null; it does not calculate how large either practical cost is.
Worked Example: Interpreting a Mean Measurement Near a Legal Threshold
In a fictional regulatory review, a court-appointed technical team measures a contaminant at randomly selected sites. The legal threshold is 50 units. Let \(\mu\) be the true mean contaminant level, in units, across the population of sites defined for the review. The team tests \(H_0:\mu=50\) against \(H_a:\mu>50\). For a random sample of 10 sites, \(\bar{x}=51.2\) units and \(s=4.0\) units. Use \(\alpha=0.05\) and consider what the two errors could mean.
State. A Type I error would be concluding that the population mean exceeds 50 units when the null claim is true. A Type II error would be failing to find convincing evidence that the mean exceeds 50 units when the null claim is false.
Plan and conditions. A one-sample t test is appropriate for a quantitative response compared with a fixed benchmark. Assume the sites were selected randomly from the population at issue and that the measurements are independent. If sites were sampled without replacement, 10 should be less than 10% of that population. Since \(n=10\) is small, the sample measurements should show no strong skewness or outliers. A real analysis would check the sample plot and the sampling process; these conditions cannot be confirmed from the summary statistics alone.
Do. The standard error and test statistic are:
The upper-tail p-value is approximately 0.184, rounded. Since \(0.184>0.05\), fail to reject \(H_0\). The data do not provide convincing evidence that the population mean contaminant level exceeds 50 units.
A 95% confidence interval uses \(t^*=2.262\) with 9 degrees of freedom. Its margin of error is \(2.262(1.2649)\approx2.8612\) units, giving \(51.2\pm2.8612\), or approximately \((48.34,54.06)\) units, rounded. The interval includes values below and above the threshold, so it does not pin down which side of 50 the population mean lies on.
Conclude and weigh the costs. If the true mean were 50 units, concluding that it exceeds 50 would be a Type I error. Depending on the rules and evidence in the case, a false conclusion could contribute to an unjustified order or penalty. If the true mean were above 50, failing to reject could be a Type II error; a real exceedance might not be acted on, potentially leaving people exposed. This statistical analysis concerns a population mean. It does not by itself establish the level at every site, identify responsibility, or determine a legal verdict. Decision-makers must apply the relevant legal standards and consider the broader evidence.
Common Mistakes and AP Exam Tips
- Calling an error “more likely” because it seems more serious: Severity and probability are different questions. The practical consequences may make one error less acceptable, but that does not show how likely either error is in a particular study.
- Changing the definition of Type I or Type II: A false alarm is a Type I error only when it means rejecting a true null hypothesis. Missing a real effect is Type II only when the null is false and the test fails to reject it.
- Claiming that an error definitely occurred: The population mean is generally unknown. Use conditional language: “If the true mean were ..., then this decision would be a Type ... error.”
- Treating the p-value as a measure of practical cost: A p-value describes evidence against \(H_0\), under the null model. It does not count dollars, injuries, unfair outcomes, or lost production.
- Assuming an interval settles the decision: A confidence interval shows plausible values for the population mean under the procedure. A wide interval that crosses a practical threshold can leave important uncertainty unresolved.
- Ignoring who bears the consequences: In a quality-control example, a factory might pay for downtime while customers bear risks from underfilled products. Name the affected groups when comparing costs in context.
- Overstating what a mean test proves: A conclusion about a population mean is not a conclusion about every individual measurement. In legal settings especially, mean evidence is only one possible part of a broader decision.
For a strong AP response, state the relevant population mean and units, connect the error to the correct test decision and population truth, and explain a plausible consequence in context. Keep the statistical conclusion separate from the practical judgment. If the question asks which error is worse, explain what harm might follow and to whom; do not claim that the sample determines the answer.
Check Your Understanding
For each question, keep the statistical error distinct from its possible practical consequence.
- A clinic tests whether a treatment lowers mean symptom duration below 6 days. Describe a possible Type I error and one consequence it could have.
- A factory tests \(H_0:\mu=250\) against \(H_a:\mu<250\), where \(\mu\) is mean package weight in grams. What would a Type II error mean, and who might be affected?
- A sample has \(\bar{x}=52\), \(s=5\), and \(n=25\). What standard error would be used for a one-sample t test about a population mean? Show the calculation.
- Why does a confidence interval that crosses a legal threshold not, by itself, determine which legal action should be taken?
- In a cost comparison, why should you discuss who bears the consequences as well as how severe those consequences might be?