Tutorials › AP Statistics › Choosing Alpha Based on Error Consequences

Inference decisions and errors · Tutorial 587 of 1000

Choosing Alpha Based on Error Consequences

Use the relative costs of Type I and Type II errors to choose a significance level, while recognizing the tradeoff and fixing alpha before seeing the results.

Intermediate 10 min read

What You'll Learn

  • Explain how alpha controls the chance of a Type I error when the null hypothesis is true.
  • Compare alpha = 0.01 and alpha = 0.10 in terms of strictness and false-alarm risk.
  • Use the relative consequences of Type I and Type II errors to justify a choice.
  • Describe the tradeoff between a lower alpha and the chance of detecting a real effect.
  • Explain why alpha must be selected before examining the data.
  • Distinguish a significance level from the probability that a particular conclusion is wrong.

Let the Consequences Guide the Significance Level

In “Consequences of Type I Versus Type II Errors,” you learned to translate each error into what it would mean in a real setting. This tutorial uses that comparison to make a related decision: should a test use \(\alpha=0.01\) or \(\alpha=0.10\)?

The significance level \(\alpha\) is the probability of a Type I error when \(H_0\) is true, as introduced in “Significance Level as the Probability of a Type I Error.” Choosing \(\alpha=0.01\) sets a more stringent standard for rejecting \(H_0\) than choosing \(\alpha=0.10\). The smaller significance level makes a false positive less likely under the null model, but it also makes it harder for the test to reject \(H_0\) when a real effect exists.

Key idea: Choose alpha by considering the consequences of both kinds of error. A smaller alpha is often appropriate when a Type I error would be especially serious. A larger alpha may be reasonable when missing a real effect is especially costly and a false alarm has manageable consequences. Neither value is best in every setting.

This is a decision made when planning a test, before looking at the sample results. It is not a choice to make afterward to obtain a preferred conclusion. The hypotheses define the question; alpha sets how strong the evidence must be for rejecting \(H_0\). The test’s p-value is then compared with that preselected alpha, as in “Making a Decision in a Chi-Square Test.”

What Changes When Alpha Changes?

Suppose the same data produce a p-value of \(0.04\). If the planned significance level is \(0.10\), then \(0.04\le0.10\), so the decision is to reject \(H_0\). If the planned significance level is \(0.01\), then \(0.04>0.01\), so the decision is to fail to reject \(H_0\). The data and p-value have not changed; the preselected evidence standard has.

A lower alpha means fewer results will lead to rejection of \(H_0\). That reduces the Type I error probability under the null model. For the same sample size and the same specified real effect, the stricter rejection standard also makes the test less likely to detect that effect. In the language of the earlier tutorial on Type II errors, the chance of a Type II error can increase, and power can decrease. The next tutorial develops the probability of a Type II error, called \(\beta\), in more detail.

Compare the choices:
  • \(\alpha=0.01\): A more demanding standard for rejecting \(H_0\); lower Type I error risk, with a greater chance of failing to detect a real effect for a fixed sample size and effect.
  • \(\alpha=0.10\): A less demanding standard for rejecting \(H_0\); higher Type I error risk, with a greater chance of detecting a real effect for a fixed sample size and effect.

The practical question is not simply “Which alpha is more cautious?” The question is “Cautious about which mistake?” A strict standard guards more against a false alarm. A more permissive standard can make it easier to act on a signal of a real effect. Weigh the consequences of both mistakes, and consider whether other safeguards—such as follow-up testing or a reversible pilot—can reduce the harm of either error.

A Context-Based Decision Process

Use the following reasoning to justify a choice between \(0.01\) and \(0.10\). Start with the meaning of the hypotheses and decisions, rather than treating the numbers as universal rules.

1
Translate both errors.
Describe the Type I error and Type II error as outcomes in the situation.
2
Compare their consequences.
Consider who could be affected, how serious the harm could be, and whether it can be undone or limited.
3
Choose the priority.
If avoiding a false alarm is the priority, justify the stricter \(\alpha=0.01\). If avoiding a missed effect is the priority and a false alarm is tolerable, \(\alpha=0.10\) may be defensible.
4
State the tradeoff and decide in advance.
Explain what risk the choice reduces and what risk it may increase. Set alpha before examining the sample results.

This process does not calculate the probability or cost of every possible outcome. It makes the reasoning behind the chosen error standard explicit. If the consequences are uncertain or depend on whose perspective is considered, say so rather than presenting a value judgment as a mathematical fact.

Worked Examples: Choosing Between 0.01 and 0.10

Worked Example: Monitoring a Community Water Supply

A public health team tests whether the proportion of water samples from a community that exceed a safety threshold is greater than the established background proportion. The null hypothesis represents no increase above that background level; the alternative represents an increase. Rejecting \(H_0\) would prompt a costly temporary shutdown and investigation. Failing to reject \(H_0\) when the unsafe proportion really has increased could leave residents exposed. Which alpha is more defensible?

Translate the errors: A Type I error would mean concluding that the unsafe-sample proportion has increased when it has not, leading to an unnecessary shutdown and investigation. A Type II error would mean failing to detect a real increase, potentially allowing exposure to continue.

Compare the consequences: The possible health effects of an undetected increase could be serious, but a false alarm also has substantial costs. The team should consider the severity and reversibility of those harms, how quickly follow-up testing can be done, and whether another safety procedure acts when there is a warning sign.

Choose and qualify: If the team judges that a false alarm would trigger major disruption and that independent checks can quickly identify a real hazard, it could justify \(\alpha=0.01\) to limit the chance of an unnecessary shutdown. But if the health risk from even a short delay is severe and shutdown is a manageable precaution, the consequences may support \(\alpha=0.10\), making it easier to act on evidence of an increase. There is no automatic answer from the setting alone; the decision depends on the stated risks and safeguards.

Worked Example: Piloting a Low-Cost Tutoring Reminder

A school tests whether sending families a new text reminder increases the proportion of students who attend scheduled tutoring sessions. If \(H_0\) is true, the reminder does not increase the population proportion. The school could expand the reminder program if it rejects \(H_0\). The messages are inexpensive, and a small pilot could be discontinued if it proves unhelpful. Which significance level could make sense?

Translate the errors: A Type I error would mean concluding that the reminder increases attendance when it does not, leading the school to spend time and money on a program without a real attendance benefit. A Type II error would mean failing to find evidence of an increase when the reminder really does help students attend.

Compare the consequences: Under the scenario’s assumptions, the program is inexpensive and reversible, so acting on a false positive during a limited pilot may have modest consequences. Missing a useful reminder program could leave students without support. These considerations can make the less stringent \(\alpha=0.10\) a defensible choice for an exploratory pilot.

State the tradeoff: At \(\alpha=0.10\), the school accepts a higher Type I error probability under \(H_0\) than it would at \(\alpha=0.01\). The choice does not prove that the reminder works, and it does not guarantee that a real increase will be detected. It reflects the judgment that a false alarm in a limited, low-cost pilot is tolerable relative to the possibility of overlooking a helpful program. If the school were deciding whether to commit a large budget permanently, the consequences of a false positive might be greater and a stricter standard could be more appropriate.

Worked Example: A p-Value Near the Two Candidate Levels

A research team tests whether a new appointment system increases the proportion of patients who arrive on time. Before collecting data, the team chooses \(\alpha=0.01\), because it wants strong evidence before paying for a costly system-wide change. The test later produces a p-value of \(0.04\). What is the decision, and what would change if the team had selected \(\alpha=0.10\) in advance?

Use the planned standard: With \(\alpha=0.01\), compare \(0.04\) with \(0.01\). Because \(0.04>0.01\), the team fails to reject \(H_0\). In context, the results do not provide convincing evidence at the 1% significance level that the new system increases the proportion of patients arriving on time.

Compare the alternative plan: If the team had selected \(\alpha=0.10\) before seeing the data, then \(0.04\le0.10\), so it would reject \(H_0\). At the 10% significance level, the results would provide convincing evidence that the system increases the proportion arriving on time. The different decisions arise because the standards differ, not because the sample changed.

Explain why the choice matters: The team cannot inspect the p-value, notice that it falls between the two levels, and then choose \(\alpha=0.10\) to claim a result it prefers. Doing that makes the significance level depend on the observed data and undermines its planned role in limiting Type I errors. The team should report the result using the alpha it selected in advance and describe its decision accurately.

Common Mistakes and AP Exam Tip

  • Treating \(0.01\) as always better: It offers a lower Type I error probability under \(H_0\), but can make a real effect harder to detect. Explain which error the stricter choice is meant to guard against.
  • Treating \(0.10\) as proof of a real effect: A larger alpha does not make the effect more likely to exist. It makes rejection of \(H_0\) possible for a wider range of p-values, while accepting greater Type I error risk.
  • Choosing alpha after seeing the p-value: This changes the standard in response to the result. Full-credit reasoning says alpha is selected before data collection or before examining the results.
  • Calling alpha the probability that \(H_0\) is true: Alpha is a Type I error probability when \(H_0\) is true; it is not the probability that the null hypothesis is true or that a particular conclusion is wrong.
  • Ignoring Type II error consequences: A low alpha reduces false-alarm risk but can make a missed real effect more likely for fixed conditions. Discuss both kinds of error, even when one is the main concern.
  • Giving a choice without a reason: “Use 0.01 because it is stricter” is incomplete. A stronger answer translates the errors, identifies the more serious consequence under the scenario’s assumptions, and connects that priority to the alpha choice.
Key takeaway: Choose between \(\alpha=0.01\) and \(\alpha=0.10\) by weighing the consequences of a false positive against those of a missed effect. A lower alpha reduces Type I error risk but can reduce the chance of detecting a real effect. State the tradeoff, justify the priority in context, and select alpha before examining the data.

Check Your Understanding

For each question, explain how the consequences of the errors relate to the choice of significance level.

  1. A company tests whether a new safety feature reduces the proportion of devices that overheat. Describe what a Type I error and a Type II error would mean.
  2. In the safety-feature setting, what additional information could help determine whether \(\alpha=0.01\) or \(\alpha=0.10\) is more defensible?
  3. A researcher selects \(\alpha=0.01\) before a study. The p-value is \(0.06\). State the decision and explain why changing to \(\alpha=0.10\) after seeing the p-value is not appropriate.
  4. Explain why choosing \(\alpha=0.10\) does not mean there is a 10% chance that a particular conclusion is wrong.
  5. Give one situation in which \(\alpha=0.10\) could be reasonable, and state which error the decision-maker is especially concerned about.