Tutorials › AP Statistics › Common Misconceptions About Errors and Significance

Inference errors and practical significance · Tutorial 798 of 1000

Common Misconceptions About Errors and Significance

Learn to distinguish the long-run Type I error rate from the probability of a hypothesis being true, and interpret significant and nonsignificant results carefully.

Intermediate 10 min read

What You'll Learn

  • Explain what alpha means in repeated tests when the null hypothesis is true
  • Distinguish alpha from a p-value and from the probability that a hypothesis is true
  • Interpret a significant result without claiming it proves the alternative hypothesis
  • Explain why failing to reject the null hypothesis does not establish that it is true
  • Separate statistical significance from the size or practical importance of a mean difference
  • Write careful, context-based statements about test results

What Statistical Significance Does—and Does Not—Mean

In “Multiple Testing and False Positives,” you saw how significance levels describe the chance of a Type I error across repeated tests. This tutorial focuses on a common interpretation mistake: treating \(\alpha\) as the probability that the null hypothesis, \(H_0\), is true. Alpha does not tell us that. It describes the long-run behavior of a testing procedure when \(H_0\) is true.

A significant result is also easy to overstate. It does not prove that \(H_a\) is true, give the probability that \(H_0\) is true, or show that an effect is large or important. A nonsignificant result does not prove that \(H_0\) is true either. To interpret a result responsibly, keep the hypothesis, the test procedure, and the evidence from the data distinct.

Definition: The significance level \(\alpha\) is the long-run probability of a Type I error when \(H_0\) is true. A Type I error is rejecting a true null hypothesis. Alpha describes the behavior of the testing procedure under that condition; it is not the probability that \(H_0\) is true.

Alpha Is Not a Probability Assigned to a Hypothesis

Suppose a test uses \(\alpha=0.05\). This means that, if the null hypothesis is true and the testing procedure’s conditions and assumptions hold, the procedure has a 0.05 long-run probability of rejecting \(H_0\). If the same procedure were used repeatedly in many comparable situations where \(H_0\) was true, about 5% of those tests would be expected to reject it.

That statement starts with the condition “if \(H_0\) is true.” It does not reverse the condition to say “given that the test rejected \(H_0\), there is a 5% chance that \(H_0\) was true.” Those are different probability questions. Alpha alone does not answer the second one.

A hypothesis test does not assign a probability to \(H_0\) or \(H_a\). Instead, it evaluates how compatible the observed data are with \(H_0\), using a test statistic and a probability model. The hypothesis is either true or false for the population being studied, even though we may not know which. The test provides evidence for a decision; it does not reveal certainty about the hypothesis.

Alpha and the P-Value Answer Different Questions

Alpha is selected as a decision threshold for the test. The p-value is calculated from the observed data: assuming \(H_0\) is true and the test’s model and conditions hold, it is the probability of getting a test statistic at least as extreme as the one observed, in the direction or directions specified by \(H_a\). Compare the p-value with \(\alpha\) to make the reject-or-fail-to-reject decision, as you practiced in “Writing Complete Conclusions for Mean Inference Problems.”

Key distinction: Alpha is a chosen long-run Type I error rate for a testing procedure under a true \(H_0\). The p-value describes how unusual the observed result, or one more extreme, would be if \(H_0\) were true and the test assumptions held. Neither is the probability that \(H_0\) is true.

For a test using \(\alpha=0.05\), a p-value of 0.03 leads to rejecting \(H_0\), while a p-value of 0.08 leads to failing to reject \(H_0\). The p-value does not change the meaning of alpha, and alpha does not change what the p-value represents.

A p-value can be small without being the probability that the result is “due to chance.” That phrase is too vague: the p-value is calculated under a particular null hypothesis and test model, not under a general explanation called chance. It also does not tell us the chance that the result will be replicated in a future study.

Worked Example: Interpreting Alpha and a Significant Result

A fictional school researcher tests \(H_0:\mu=50\) against \(H_a:\mu>50\), where \(\mu\) is the population mean score after a new practice program. The test uses \(\alpha=0.05\) and produces a p-value of 0.032. A colleague says, “There is a 5% chance the null hypothesis is true, so we can reject it.” Evaluate the statement and report the result correctly.

State. Identify what the test result says about the population mean and clarify the meaning of \(\alpha\).

Plan. Compare the p-value with \(\alpha\) for the test decision. Then interpret alpha as a long-run Type I error rate under a true null, not as a probability that the null is true.

Do. Since \(0.032<0.05\), reject \(H_0\). The value 0.05 is the chosen significance level. If the null claim \(\mu=50\) is true and the test conditions hold, this procedure has a 0.05 long-run probability of rejecting that true claim. The p-value 0.032 means that, if \(H_0\) is true and the test model applies, the probability of a test statistic at least as large as the observed one is 0.032.

Conclude. The colleague’s interpretation of alpha is incorrect. The test provides convincing evidence that the population mean score after the program is greater than 50. It does not establish that \(H_a\) is certainly true or give the probability that \(H_0\) is true.

A Significant Result Is Evidence, Not Proof

Rejecting \(H_0\) at a chosen significance level means the data provide enough evidence, by the test’s decision rule, to reject the null claim. In AP Statistics, a careful conclusion says that the data provide convincing evidence for the alternative claim when the decision is to reject. It does not say the test has proved the alternative.

A Type I error remains possible: the test might reject \(H_0\) even though \(H_0\) is true. Alpha describes the long-run chance of that error under a true null and the stated procedure. It cannot tell us whether a particular rejection is a Type I error. As “Describing Errors in Context” emphasizes, an error can be identified only by comparing the test decision with the actual, usually unknown, state of the population.

Statistical significance also does not measure the size of a difference. A small mean difference can produce a small p-value, especially with a large sample, while a larger difference can fail to be statistically significant when the estimate is imprecise. As discussed in “Is a Statistically Significant Difference in Means Practically Important?”, the practical importance of a result depends on its size and context, not just on whether \(p\leq\alpha\).

Worked Example: A P-Value Is Not the Probability the Null Is True

In a fictional gardening experiment, a test compares the mean growth of seedlings receiving a nutrient mix with a benchmark mean of 12 centimeters. The hypotheses are \(H_0:\mu=12\) and \(H_a:\mu\ne12\). A two-sided test reports \(p=0.018\) at \(\alpha=0.05\). One student says, “There is a 98.2% chance the nutrient mix changes mean growth.” Is that interpretation justified?

State. Decide whether the data provide evidence of a difference from 12 centimeters and assess what the p-value means.

Plan. Compare \(p\) with \(\alpha\). Interpret the p-value conditionally on \(H_0\) and the test assumptions, and avoid treating it as a probability that either hypothesis is true.

Do. Since \(0.018<0.05\), reject \(H_0\). Under the assumption that the population mean is 12 centimeters and the test conditions hold, the probability of a test statistic at least as extreme as the observed one in either direction is 0.018. The p-value is not \(1-0.018=0.982\) as a probability that \(H_a\) is true.

Conclude. The data provide convincing evidence that the population mean growth differs from 12 centimeters. The student’s 98.2% interpretation is not justified: the p-value does not provide the probability that the nutrient mix changes the mean. The test also does not, by itself, establish that the mix caused any difference; that depends on the study design.

Failing to Reject Does Not Prove the Null

When \(p>\alpha\), the test fails to reject \(H_0\). That decision means the data do not provide sufficiently strong evidence against the null claim according to the chosen procedure. It does not mean that the null has been proved, accepted as true, or assigned a high probability.

A test may fail to reject a false null hypothesis. This is a Type II error, as explained in “Type II Error in a Test About Means.” Such an outcome can occur when the sample is small, variability is high, or the actual difference is modest. A large p-value therefore does not show that there is no difference; it shows that the data did not provide convincing evidence of a difference using that test.

If the research question is whether an effect is small enough to be practically negligible, a test that fails to reject a point null is not enough to establish that conclusion. The question may call for an estimate and a confidence interval considered against a practical-importance threshold, as in “Using Confidence Intervals to Judge Practical Importance.”

Worked Example: A Nonsignificant Result Does Not Confirm No Change

A fictional community clinic tests whether a new appointment reminder changes the population mean waiting time from 30 minutes. The test uses \(H_0:\mu=30\) and \(H_a:\mu\ne30\), with \(\alpha=0.05\). The reported p-value is 0.18. A manager concludes, “The reminders do not change waiting time because the null hypothesis is true.” Assess the conclusion.

State. Determine the test decision and describe what the data support about the population mean waiting time.

Plan. Compare the p-value with alpha. If the test fails to reject, report insufficient evidence for the alternative rather than claiming that the null is true.

Do. Because \(0.18>0.05\), fail to reject \(H_0\). The result does not meet the test’s threshold for evidence that the mean waiting time differs from 30 minutes. It does not demonstrate that \(\mu=30\), and it does not assign a probability to that equality.

Conclude. The data do not provide convincing evidence that the population mean waiting time differs from 30 minutes. The manager’s claim that the null is true goes beyond what this test establishes. A Type II error is possible if the true mean differs from 30 minutes.

Common Mistakes and AP Exam Tips

  • Calling alpha the probability that \(H_0\) is true: Say instead, “If \(H_0\) is true, the procedure has a long-run Type I error probability of \(\alpha\).” The condition matters.
  • Calling the p-value the probability that \(H_0\) is true: A p-value is calculated assuming \(H_0\) is true. It describes the probability of results at least as extreme as those observed under that assumption and the test model.
  • Claiming a significant result proves \(H_a\): A significant result provides evidence for the alternative, but a Type I error is possible. Use “provides convincing evidence,” not “proves.”
  • Accepting \(H_0\) after a nonsignificant result: State “fail to reject \(H_0\)” and explain that the data do not provide convincing evidence for \(H_a\). Do not claim that no effect exists.
  • Confusing statistical significance with importance: A small p-value is not a measure of the size or practical value of a difference. Describe the estimated difference in context and original units when available.
  • Ignoring the study design: A test may show evidence of an association or difference, but a cause-and-effect conclusion requires an appropriate randomized experiment. Significance alone does not establish causation.

For full credit, state the decision, connect it to the population parameter and claim, and interpret the evidence in context. Keep the direction of reasoning clear: alpha describes a procedure assuming a true null; the p-value describes the data assuming a true null; neither gives the probability that the null is true.

Key takeaway: Alpha is the long-run Type I error rate when \(H_0\) is true, not the probability that \(H_0\) is true. A significant result provides evidence against \(H_0\), not proof of \(H_a\); failing to reject \(H_0\) does not prove it true.

Check Your Understanding

Answer each question by distinguishing what the test procedure controls from what the data provide evidence about.

  1. In a mean test with \(\alpha=0.01\), what does the 0.01 describe when \(H_0\) is true?
  2. A test has \(p=0.04\) and uses \(\alpha=0.05\). What is the decision, and why is it incorrect to say that \(H_0\) has a 4% chance of being true?
  3. A student says, “Because \(p=0.12\), there is a 12% chance that the alternative hypothesis is true.” Explain the mistake.
  4. What is an appropriate contextual conclusion when \(p>\alpha\)? Why should it not say that the null hypothesis has been proved?
  5. Why does statistical significance alone not show that a mean difference is practically important?