Tutorials › AP Statistics › Significance Level as the Probability of a Type I Error

Inference decisions and errors · Tutorial 585 of 1000

Significance Level as the Probability of a Type I Error

See how alpha controls the long-run probability of rejecting a true null hypothesis, and learn to explain what alpha = 0.05 means in context.

Intermediate 10 min read

What You'll Learn

  • Express the significance level as \(P(\text{reject }H_0\mid H_0\text{ is true})\)
  • Interpret alpha = 0.05 as a long-run probability for a testing procedure
  • Explain what a 5% Type I error rate means in a proportion-testing context
  • Distinguish alpha from a p-value and from the probability that a particular conclusion is wrong
  • Describe why a chosen alpha is a threshold, not a guarantee of an exact number of errors

From Type I Error to Significance Level

In “Two Possible Errors in a Significance Test” and the tutorials that followed, you learned that a Type I error is rejecting a true null hypothesis. The significance level, written \(\alpha\), connects that definition to the chance of making this error. It describes how often a testing procedure is designed to reject \(H_0\) when \(H_0\) is true.

The vertical bar in \(P(\text{reject }H_0\mid H_0\text{ is true})\) means “given that.” So this notation asks: if the null hypothesis is in fact true, what is the probability that the test procedure will reject it? The answer is the probability of a Type I error.

Definition: The significance level \(\alpha\) is the probability of a Type I error for a significance-testing procedure. In probability notation, \(\alpha=P(\text{reject }H_0\mid H_0\text{ is true})\). For a composite null hypothesis, the procedure is designed to keep this probability at or below \(\alpha\) for null values; it is often equal to \(\alpha\) at the boundary value.

A significance test uses sample data to decide whether to reject \(H_0\) or fail to reject it. Before looking at the data, the researcher chooses \(\alpha\), such as \(0.05\). The test is then set up so that, when the null model is true, the chance of getting a result that leads to rejection is controlled at the chosen level.

This is a long-run description of the procedure, not a prediction that a particular test will make a fraction of an error. For any one completed test, the decision either is or is not a Type I error; we usually do not know the actual truth with certainty. Alpha tells us how often the procedure would make that error over repeated uses when the null hypothesis is true.

What Does Alpha = 0.05 Mean?

Suppose a study tests whether a new community program increases the population proportion of residents who sort food scraps for composting. Let \(p\) be that population proportion after the program is introduced. The hypotheses are \(H_0:p=0.30\) and \(H_a:p>0.30\), and the test uses \(\alpha=0.05\).

If \(H_0\) is true, the population proportion is \(0.30\). In repeated applications of this same testing procedure under that null condition, the probability of rejecting \(H_0\) is \(0.05\). Those rejections would be Type I errors: the procedure would indicate convincing evidence that the proportion exceeds \(0.30\), although it does not.

Interpretation: With \(\alpha=0.05\), if the null hypothesis is true, this testing procedure has a 5% probability of rejecting it and making a Type I error. Equivalently, in many repeated uses when the null model is true, about 5% of the tests would reject \(H_0\) just by chance.

“About 5%” is a long-run rate, not a promise about a particular batch of tests. If the procedure is repeated 100 times while the null is true, it does not have to produce exactly five Type I errors. A particular set of repetitions could produce fewer or more.

For a null hypothesis that covers a range of values, such as \(H_0:p\le 0.30\), the interpretation needs a small qualification. The test is designed to limit the probability of a Type I error to no more than \(0.05\) for values in the null range. The probability is often exactly \(0.05\) at the boundary, \(p=0.30\), and smaller for some values farther inside the null range. At the AP Statistics level, describe alpha as the procedure’s Type I error rate under the null model, and identify the relevant null value when the setting makes it important.

Worked Examples: Interpreting Alpha in Context

Worked Example: A Composting Program

A city evaluates whether a composting program increases the proportion of residents who sort food scraps. The hypotheses are \(H_0:p=0.30\) versus \(H_a:p>0.30\), and the test uses \(\alpha=0.05\). Interpret the significance level.

Identify the Type I error: A Type I error would occur if the test rejected \(H_0\) and concluded that the program increases the population proportion above \(0.30\), even though the true population proportion is \(0.30\).

Interpret alpha: If the true proportion is \(0.30\), the probability that this testing procedure rejects \(H_0\) is \(0.05\). In repeated uses of the procedure under this null condition, about \(5\%\) of the tests would make this false-increase conclusion.

Check the scope: This does not say there is a \(5\%\) probability that the null hypothesis is true. It also does not say that this particular rejection, if one occurs, has a \(5\%\) chance of being wrong. Alpha describes the procedure’s long-run behavior when the null is true.

Worked Example: Counting False Alarms in Repeated Tests

A water-quality team tests whether a stream’s population proportion of samples exceeding a specified algae level is greater than \(0.08\). Assume the null value \(p=0.08\) is true. The team uses a procedure with \(\alpha=0.05\) for 600 separate, repeated tests conducted under the same conditions. About how many Type I errors would the procedure produce in the long run?

Use the long-run rate: When the null is true, the probability of a Type I error on each use is \(0.05\). The expected number in 600 uses is the number of uses multiplied by that probability:

$$ 600(0.05)=30 $$

Interpret the result: The long-run expected number is 30 Type I errors. In context, if the stream’s true proportion is \(0.08\), the procedure would falsely indicate that the proportion exceeds \(0.08\) about 30 times per 600 repeated tests in the long run.

State what is not guaranteed: Thirty is an expected count, not a fixed quota. One set of 600 tests might produce 26 Type I errors, another might produce 34, and other outcomes are possible. The calculation uses the rate in the test design; it does not predict the exact result of one group of repetitions.

Worked Example: An Observed P-Value and Two Alpha Levels

A school district tests whether an online reminder increases the population proportion of families who complete a registration form on time. The null hypothesis sets the proportion at \(0.70\), and the alternative is that it is greater than \(0.70\). A calculator reports a p-value of \(0.032\). Compare the decision at \(\alpha=0.05\) and \(\alpha=0.01\), then interpret each significance level.

At \(\alpha=0.05\): Since \(0.032\le 0.05\), the district rejects \(H_0\). If \(H_0\) is true, this testing procedure has a \(5\%\) Type I error probability: it may reject the true null and indicate an increase when the true proportion is \(0.70\).

At \(\alpha=0.01\): Since \(0.032>0.01\), the district fails to reject \(H_0\). A procedure using this significance level is designed to have a Type I error probability of \(1\%\) under the null model (or no more than \(1\%\) across a composite null).

Explain the comparison: The p-value \(0.032\) is evidence measured from the sample and evaluated assuming the null model. Alpha is the decision threshold chosen for the procedure and its Type I error rate. The p-value is not the probability that the null is true, nor is it the probability that the district’s conclusion is wrong.

Alpha Is Not the Probability a Conclusion Is Wrong

The conditional wording matters. Alpha is \(P(\text{reject }H_0\mid H_0\text{ is true})\): it starts by assuming the null is true and asks how often the procedure rejects. It is not \(P(H_0\text{ is true}\mid\text{the test rejected})\). These probabilities condition on different information and are not interchangeable.

After a test rejects \(H_0\), alpha alone does not tell us the probability that this particular rejection is a Type I error. To make that probability claim, we would need information beyond the significance level, including how plausible the null was before the data and how well the procedure detects alternatives. AP test conclusions instead describe the evidence: for example, whether the sample provides convincing evidence for the alternative in context.

Alpha is also different from the p-value. Alpha is selected as a threshold before the test decision; the p-value is calculated from the observed data. The decision rule compares them: reject \(H_0\) when the p-value is less than or equal to \(\alpha\), and fail to reject \(H_0\) when it is greater than \(\alpha\). As the example shows, the same p-value can lead to different decisions under different preselected significance levels.

Key distinction: Alpha is the test procedure’s Type I error probability under the null model. A p-value is the probability, assuming \(H_0\) is true, of obtaining a test statistic at least as extreme as the one observed in the direction of \(H_a\). Neither quantity is the probability that \(H_0\) is true after seeing the data.

How to State Alpha Clearly

A strong interpretation names the null condition, the test decision that would be an error, and the population outcome being studied. Use conditional language: “If the true population proportion is …, the probability this procedure rejects \(H_0\) is …” Avoid describing alpha as a percentage of individuals, a probability that a hypothesis is true, or a guaranteed number of mistakes.

1
Name the null condition.
State what is assumed to be true, such as a population proportion equal to the claimed value.
2
Name the Type I error.
Describe the incorrect conclusion the test would reach by rejecting that true null hypothesis.
3
Interpret alpha as a long-run probability.
Say how often the procedure would make that error over repeated uses when the null model is true.

For the registration example, a complete statement is: “If the true proportion of families completing registration on time is \(0.70\), this procedure has a \(5\%\) probability of rejecting the null and concluding that the reminder increases the proportion.” This identifies the condition, the error, the probability, and the context.

Common Mistakes and AP Exam Tip

  • Reversing the conditional probability: Alpha is the probability of rejecting \(H_0\) given that \(H_0\) is true. It is not the probability that \(H_0\) is true given that it was rejected.
  • Calling alpha the probability that a specific conclusion is wrong: Alpha describes repeated use of a procedure under the null condition. It does not tell us whether one observed rejection is an error.
  • Confusing alpha with the p-value: Alpha is selected as the significance threshold; the p-value comes from the sample. State which number is which before comparing them.
  • Claiming an exact error count: An alpha of \(0.05\) does not guarantee exactly five Type I errors in every 100 tests. Say “about 5% in the long run” when the null is true.
  • Leaving out the context: “There is a 5% chance of error” is vague. Full-credit wording identifies the false conclusion and the null condition, such as “If the true proportion is \(0.70\), there is a 5% chance the test will conclude it is greater than \(0.70\).”
  • Overstating alpha for a composite null: When \(H_0\) covers a range, the procedure controls the Type I error probability at or below alpha across that range; it need not equal alpha at every null value.
Key takeaway: The significance level is the probability of rejecting a true null hypothesis: \(\alpha=P(\text{reject }H_0\mid H_0\text{ is true})\). Thus, \(\alpha=0.05\) means a 5% long-run Type I error probability under the null model—not a 5% probability that a particular conclusion is wrong.

Check Your Understanding

Answer each question using the conditional, long-run meaning of the significance level.

  1. A test checks whether a population proportion exceeds \(0.12\), using \(\alpha=0.05\). In context, explain what the 5% refers to if the true proportion is \(0.12\).
  2. A researcher says, “My test used \(\alpha=0.01\), so there is a 1% chance the null hypothesis is true.” Explain the error in this statement.
  3. A testing procedure has \(\alpha=0.10\) and is used 250 times under conditions where each null hypothesis is true. Calculate the long-run expected number of Type I errors, and explain why that number is not guaranteed.
  4. A test produces a p-value of \(0.04\). Explain why \(0.04\) is not automatically the probability that the test’s conclusion is wrong.
  5. Write a complete, in-context interpretation of \(\alpha=0.05\) for a test of whether a recycling campaign increases the population proportion of households that recycle.