Tutorials › AP Statistics › Writing About Errors and Practical Significance on the Exam

Inference errors and practical significance · Tutorial 800 of 1000

Writing About Errors and Practical Significance on the Exam

Practice connecting a test’s decision to the error it could represent and judging whether an estimated mean effect matters in practice.

Intermediate 10 min read

What You'll Learn

  • Describe a Type I error by stating what was wrongly concluded when the null hypothesis is true.
  • Describe a Type II error by stating what was missed when the null hypothesis is false.
  • Explain a possible real-world consequence without treating an error as certain.
  • Separate a test’s statistical decision from the practical importance of an effect.
  • Compare an estimate and confidence interval with a practical-importance threshold in context.
  • Write careful conclusions that avoid claiming more than the data support.

Connect the Decision to What Could Go Wrong

A test result can lead someone to change a policy, adjust a process, or decide that no action is needed. The decision may be reasonable given the data, yet still be wrong because the null hypothesis is either true or false in reality. To describe an error well, connect the test’s decision to the population claim and then explain what could happen as a result.

In “Describing Errors in Context,” you learned the definitions of Type I and Type II errors. This tutorial focuses on turning those definitions into complete exam responses, including the possible consequences of an error. It also builds on “Evaluating Claims Based on Significant Results”: a statistically significant result does not, by itself, show that an effect is practically important.

Definition: A Type I error is rejecting a true null hypothesis. In context, the test concludes that the population mean differs in the direction described by \(H_a\), even though the null claim is true. A Type II error is failing to reject a false null hypothesis. In context, the test does not find convincing evidence for \(H_a\), even though the population claim described by \(H_a\) is true.

The error description depends on both parts of the test outcome: what the test decision was and what is actually true about the population. The hypotheses name the population claim; the decision names what the test did. The error is the mismatch between those two.

Test decisionPopulation truthWhat happened
Reject \(H_0\)\(H_0\) is trueType I error: a true null claim was rejected.
Fail to reject \(H_0\)\(H_0\) is falseType II error: the test failed to detect a real departure from the null claim.
Reject \(H_0\)\(H_0\) is falseThe decision is consistent with the population truth.
Fail to reject \(H_0\)\(H_0\) is trueThe decision is consistent with the population truth.

A test decision does not reveal which row applies, because the population truth is not known. So describe an error as a possible consequence: “If the null hypothesis is actually true, rejecting it would be a Type I error.” Do not write as though the test has shown that an error occurred.

Describe the Error and Its Consequence

A complete contextual description names the population parameter or quantity, the incorrect conclusion, and a realistic consequence. Keep the consequence conditional: it is what could happen if the corresponding error occurs. For example, a false alarm might waste money or prompt an unnecessary intervention; a missed effect might leave a problem unaddressed.

Writing pattern: If \(H_0\) is true, rejecting it would be a Type I error: we would conclude [the contextual claim in \(H_a\)] when [the contextual null claim] is actually true. This could lead to [a plausible consequence]. If \(H_0\) is false, failing to reject it would be a Type II error: we would fail to detect [the real contextual effect], potentially leading to [a plausible consequence].

State the null claim accurately. If \(H_0\) says a population mean equals a target, a Type I error is not simply “being wrong”; it is concluding that the mean differs in the specified direction when it actually equals that target. Similarly, a Type II error is not proof that there is no effect. It is the possibility that a test fails to detect an effect that is really present.

Worked Example: A False Alarm About Battery Life

A fictional repair team tests whether the mean operating time of a model of emergency radio batteries is below a 40-hour target. The hypotheses are \(H_0:\mu=40\) hours and \(H_a:\mu<40\) hours, where \(\mu\) is the population mean operating time for this battery model. The test rejects \(H_0\) at \(\alpha=0.05\). The team is considering replacing the batteries sooner.

State the possible error. Because the test rejected \(H_0\), the possible error is Type I. It occurs if the population mean operating time is actually 40 hours, despite the test’s conclusion that there is convincing evidence the mean is below 40 hours.

Describe a consequence. If that Type I error occurs, the repair team might replace batteries earlier than necessary. This could increase replacement costs and send usable batteries to disposal. These are possible consequences of the error, not outcomes the test has established.

The significance level \(\alpha=0.05\) is the long-run probability of a Type I error when \(H_0\) is true, as explained in “Alpha as the Type I Error Rate in t Tests.” It is not the probability that this particular rejection is wrong or that \(H_0\) is true.

Worked Example: A Missed Increase in Sensor Delay

A fictional monitoring team tests whether the population mean delay of a sensor exceeds 2 seconds. Here, \(\mu\) is the true mean delay for sensors in the target system. The hypotheses are \(H_0:\mu=2\) seconds and \(H_a:\mu>2\) seconds. The test fails to reject \(H_0\). A delay above 2 seconds could matter because the sensor is used to flag equipment problems.

State the possible error. Because the test failed to reject \(H_0\), the possible error is Type II. It occurs if the true population mean delay is actually greater than 2 seconds, even though the test did not provide convincing evidence of an increase.

Describe a consequence. If that Type II error occurs, the team could leave a slow sensor in service and fail to respond as quickly to equipment problems as intended. The test result does not establish that the mean delay is 2 seconds; it only says that the data did not lead to rejection of that null claim.

As “Power and Type II Error Probability” explains, the probability of a Type II error depends on a specified true alternative. Without specifying such a true mean and the test’s conditions, do not assign a numerical probability to the possibility in this example.

Practical Importance Is a Separate Judgment

A statistically significant result addresses evidence against \(H_0\), not whether an effect is large enough to matter. To judge practical importance, describe the estimated effect in the response’s original units and compare it with a meaningful, context-specific threshold. The threshold should reflect what people in that setting consider consequential, rather than being chosen solely because it makes a result look important.

An estimated effect smaller than the threshold does not automatically mean the effect is unimportant in every possible sense. Likewise, an estimate above the threshold does not erase uncertainty. “Using Confidence Intervals to Judge Practical Importance” explains why the full interval can help show which effect sizes remain plausible.

Keep the questions distinct:

  • Statistical evidence: Does the test provide convincing evidence against \(H_0\) at the stated \(\alpha\)?
  • Estimated size: What is the observed mean or mean difference, and what are its direction and units?
  • Practical judgment: How does the estimated size, and where relevant the confidence interval, compare with the practical threshold?
  • Possible error: If the decision is wrong, what could follow in this particular setting?

Worked Example: A Significant Time Saving Near the Practical Threshold

A fictional logistics company compares two packing workflows. A random sample of 36 packing stations from a system with 500 stations uses both workflows in randomized order, with enough reset time between trials to avoid carryover. For each station, define \(d\) as standard-workflow time minus revised-workflow time for a fixed task. Thus, a positive difference means the revised workflow saved time. The sample mean difference is \(\bar{d}=4.6\) minutes and the standard deviation of the differences is \(s_d=3.0\) minutes. The differences are roughly symmetric with no apparent outliers. The company considers savings of at least 5 minutes per task practically important. Test at \(\alpha=0.05\), then assess practical importance.

State. Let \(\mu_d\) be the true mean number of minutes saved per task by the revised workflow across the population of packing stations represented by the sample. The hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d>0\).

Plan. Use a paired t test because each station was measured under both workflows, and the analysis uses one difference per station. The stations were randomly sampled, and the 10% condition is met because \(36<50\), which is 10% of 500. The differences from distinct stations are treated as independent; the sampling design and 10% condition support this. The differences are roughly symmetric with no apparent outliers, supporting the t procedure for this sample size. The randomized order and reset time support a fair within-station comparison.

Do. The standard error is

$$ SE_{\bar d}=\frac{s_d}{\sqrt{n}} =\frac{3.0}{\sqrt{36}} =\frac{3.0}{6} =0.5\text{ minutes}. $$

The test statistic is

$$ t=\frac{\bar d-0}{s_d/\sqrt{n}} =\frac{4.6}{3.0/\sqrt{36}} =\frac{4.6}{0.5} =9.20. $$

The degrees of freedom are \(36-1=35\). The one-sided p-value is less than \(0.0001\). Since \(p<0.05\), reject \(H_0\).

The estimated mean saving is 4.6 minutes per task. For added context, a 95% confidence interval uses \(t^*=2.0301\) with 35 degrees of freedom:

$$ \bar d\mathbin{\pm}t^*\frac{s_d}{\sqrt{n}} =4.6\mathbin{\pm}2.0301(0.5) =4.6\mathbin{\pm}1.0151 \approx(3.585,\ 5.615)\text{ minutes}. $$

Conclude. The data provide convincing evidence that the mean time saved per task by the revised workflow is greater than zero. However, the estimate of 4.6 minutes is below the company’s 5-minute practical-importance threshold. The 95% confidence interval includes values below and above 5 minutes, so it does not establish that the true mean saving reaches that threshold. The result is statistically significant, but whether the workflow is practically worthwhile remains uncertain under the stated criterion.

If the true mean saving were zero, rejecting \(H_0\) would be a Type I error: the company could treat the revised workflow as producing a real average saving when it does not, perhaps spending time or money to adopt it unnecessarily. If the true mean saving were positive but the test had failed to reject \(H_0\), that would be a Type II error: the company could overlook a real saving. Neither error is shown to have occurred here; these are the possible consequences associated with the two types of mistaken decision.

Common Mistakes and AP Exam Tips

  • Describing an error without stating the population truth: A Type I error requires \(H_0\) to be true and rejected. A Type II error requires \(H_0\) to be false and not rejected. Include both parts in the contextual description.
  • Calling failure to reject proof of no effect: Say the test did not provide convincing evidence for \(H_a\). It does not prove \(H_0\) or show that an effect is absent.
  • Claiming an error definitely happened: The test does not reveal the unknown population truth. Use conditional wording such as “If the null claim is true, this rejection would be a Type I error.”
  • Giving a consequence that is not connected to the setting: Name a realistic result of the mistaken decision, such as unnecessary replacement or a missed delay. Make clear that it is possible, not certain.
  • Treating significance as practical importance: Give the effect estimate with direction and units, then compare it with a context-specific threshold. Do not use the p-value as a measure of how large the effect is.
  • Ignoring uncertainty around the estimate: When an interval is available, compare the whole interval with the practical threshold. An interval that crosses the threshold leaves uncertainty about whether the effect reaches it.

For full-credit communication, state the test decision accurately, describe the relevant possible error in context, and explain a plausible consequence. If practical importance is part of the question, report the estimated effect in original units and compare it with the stated threshold. As emphasized in “Writing Complete Conclusions for Mean Inference Problems,” use “convincing evidence” rather than “proof.”

Key takeaway: Describe a Type I or Type II error by matching the test decision with the population truth that would make it wrong, then explain a possible contextual consequence. Judge practical importance separately from statistical significance by comparing the effect’s size and uncertainty with a relevant threshold.

Check Your Understanding

For each question, connect the test decision to the population truth or practical threshold in context.

  1. A test rejects \(H_0:\mu=12\) against \(H_a:\mu>12\). Describe the possible Type I error in context if \(\mu\) is a mean wait time in minutes. Give one plausible consequence of that error.
  2. A test fails to reject \(H_0:\mu=0\) against \(H_a:\mu\ne0\). What population truth would make this a Type II error, and why does the decision not prove that the mean is zero?
  3. A test finds convincing evidence of a nonzero mean difference, but the estimated difference is 0.4 units and the practical threshold is 3 units. What additional information would help judge practical importance?
  4. A confidence interval for a mean saving extends from 2 to 6 minutes, while the practical threshold is 5 minutes. What does the interval suggest about whether the saving reaches the threshold?
  5. Write a conditional sentence describing a Type II error for a test about whether a treatment reduces mean recovery time, and identify one possible consequence.