Tutorials › AP Statistics › Type I Error and Multiple Tests

Inference decisions and errors · Tutorial 595 of 1000

Type I Error and Multiple Tests

See how repeated testing can raise the chance of at least one Type I error, and how a stricter per-test significance level can limit that risk.

Intermediate 10 min read

What You'll Learn

  • Define the familywise Type I error probability for a collection of tests.
  • Calculate the chance of at least one false positive in 20 independent tests when every null hypothesis is true.
  • Compare the overall false-positive chance for different numbers of tests.
  • Distinguish the expected number of false positives from the chance of getting at least one.
  • Find a per-test significance level that targets a specified overall error probability under independence.
  • Explain why the independence assumption and transparent reporting matter.

Why One Test Is Different From Many Tests

In “Significance Level as the Probability of a Type I Error,” we defined \(\alpha\) as the probability of rejecting a true null hypothesis for one test. If a researcher runs many tests, however, there are many opportunities to reject a true null. The chance of making at least one Type I error across the collection can be much larger than the significance level for any single test.

This issue can arise when a team checks many outcomes, compares many groups, or repeats analyses while looking for a result that stands out. A small p-value from one test still has its usual meaning for that test. But when many tests are considered, the chance that at least one result looks statistically significant just by chance also matters.

Definition: The familywise Type I error probability is the probability of making at least one Type I error among a collection of significance tests. A Type I error is rejecting a true null hypothesis.

The familywise probability is not automatically equal to the \(\alpha\) used for each test. Its value depends on how many null hypotheses are true, the per-test significance level, and how the test results are related. We will first study a clear case: all 20 null hypotheses are true, each test uses \(\alpha=0.05\), and the test decisions are independent.

Calculating the Chance of at Least One False Positive

For one test with a true null hypothesis, the probability of not making a Type I error is \(1-\alpha\). If the tests are independent, the probability of avoiding a Type I error on every test is the product of those probabilities. The probability of at least one Type I error is the complement.

$$ P(\text{at least one Type I error}) =1-P(\text{no Type I errors}) =1-(1-\alpha)^m $$

Here, \(m\) is the number of independent tests, all of whose null hypotheses are true, and each test uses the same significance level \(\alpha\). The independence assumption is essential to this particular formula: it lets us multiply the probabilities of avoiding an error on each test.

Worked Example: Twenty Independent Tests at the 0.05 Level

A research team plans 20 independent significance tests on different outcomes. Suppose every null hypothesis is actually true and each test uses \(\alpha=0.05\). Find the probability that the team makes at least one Type I error.

State: We want the familywise Type I error probability: the chance of rejecting at least one true null hypothesis among the 20 tests.

Plan and conditions: Use the complement of making no Type I errors. Each test has a \(0.05\) chance of a Type I error and a \(1-0.05=0.95\) chance of no Type I error. The calculation assumes all 20 null hypotheses are true, each test is conducted at the stated significance level, and the test decisions are independent.

Do: Under independence, the probability of no Type I errors is \(0.95^{20}\). Therefore,

$$ P(\text{at least one Type I error}) =1-(1-0.05)^{20} =1-0.95^{20} \approx1-0.3585 =0.6415 $$

Conclude: If all 20 null hypotheses are true and the tests are independent, there is about a \(0.6415\), or \(64.15\%\), chance of at least one false positive. Although each test has a \(5\%\) Type I error probability, the chance of at least one such error across the entire set is much higher.

Notice what this probability does not say. It does not mean that a particular significant result has a \(64.15\%\) chance of being false. It describes the chance of at least one Type I error across the collection of tests, under the stated assumptions that all null hypotheses are true and the tests are independent.

How the Number of Tests Changes the Overall Chance

With independent tests and the same per-test \(\alpha\), adding more true null hypotheses gives the team more opportunities to make a Type I error. The probability of avoiding an error on every test, \((1-\alpha)^m\), gets smaller as \(m\) increases. Its complement, the familywise Type I error probability, gets larger.

Worked Example: Comparing Five, Ten, and Twenty Tests

A team is considering running 5, 10, or 20 independent tests. Assume every null hypothesis is true and each test uses \(\alpha=0.05\). Compare the probability of at least one Type I error for the three plans.

Plan and conditions: For each plan, calculate \(1-(1-\alpha)^m\). This uses the same assumptions as the 20-test example: all null hypotheses are true, each test has the stated Type I error probability, and the test decisions are independent.

Do: For 5 tests, the probability of no Type I errors is \(0.95^5\approx0.7738\), so the chance of at least one is \(1-0.7738\approx0.2262\). For 10 tests, \(0.95^{10}\approx0.5987\), so the chance of at least one is \(1-0.5987\approx0.4013\). For 20 tests, \(0.95^{20}\approx0.3585\), so the chance of at least one is \(1-0.3585\approx0.6415\). The results are summarized below.

Number of testsProbability of no Type I errorsProbability of at least one Type I error
5\(0.95^5\approx0.7738\)\(1-0.95^5\approx0.2262\)
10\(0.95^{10}\approx0.5987\)\(1-0.95^{10}\approx0.4013\)
20\(0.95^{20}\approx0.3585\)\(1-0.95^{20}\approx0.6415\)

Conclude: Under these assumptions, the probability of at least one false positive increases from about \(0.2262\) with 5 tests to \(0.4013\) with 10 and \(0.6415\) with 20. The significance level for each individual test remains \(0.05\); it is the chance of at least one error across the whole collection that grows.

One related quantity is the expected number of Type I errors when all 20 null hypotheses are true: \(20(0.05)=1\). This is an average over many repetitions of the entire set of tests. It does not mean that exactly one false positive will occur in a particular set of 20 tests. A set could have none, one, or several.

Choosing a Per-Test Alpha for an Overall Target

A researcher may want the chance of at least one false positive across a collection to be no more than a chosen target. When the tests are independent, the formula can be rearranged to find a per-test significance level that gives that target under the assumption that all the null hypotheses are true.

Worked Example: Targeting a 0.05 Overall Chance With Twenty Tests

A team will run 20 independent tests and wants the probability of at least one Type I error to be \(0.05\), assuming all 20 null hypotheses are true. Find the per-test significance level \(\alpha\) that achieves this target.

State: We want the probability of at least one Type I error across all 20 tests to equal \(0.05\).

Plan and conditions: Use the independent-tests formula \(1-(1-\alpha)^{20}\). The calculation assumes all 20 null hypotheses are true and that the tests are independent. Solve for \(\alpha\); this is a more stringent per-test level than \(0.05\).

Do: Set the familywise probability equal to \(0.05\) and solve:

$$ 1-(1-\alpha)^{20}=0.05 $$

Thus, \((1-\alpha)^{20}=0.95\). Take the twentieth root, then subtract from 1:

$$ 1-\alpha=0.95^{1/20} \qquad \alpha=1-0.95^{1/20} \approx0.002561 $$

Conclude: A per-test significance level of about \(0.00256\) gives a \(0.05\) chance of at least one Type I error across 20 independent tests when all 20 null hypotheses are true. This illustrates the tradeoff: using a smaller \(\alpha\) reduces the chance of false positives, but, as discussed in “How Significance Level Affects Power and Type II Error,” a smaller \(\alpha\) generally reduces power when other aspects of the test stay fixed.

This calculation is specific to the stated number of tests and the independence assumption. It is not a reason to change \(\alpha\) after seeing the results. A team should decide in advance which hypotheses and outcomes it will test, how it will handle multiple tests, and how it will report the results.

Independence, Planning, and Honest Reporting

Tests are not always independent. For example, two outcomes measured on the same people may be related, so the test results may also be related. In that case, multiplying \(1-\alpha\) by itself \(m\) times does not necessarily give the exact probability of avoiding all Type I errors. The formula in this tutorial is for independent tests; without independence, the overall probability depends on how the tests are related.

The all-null assumption matters, too. The 20-test calculation describes the probability of at least one Type I error when all 20 null hypotheses are true. If some are false, rejecting those false null hypotheses is not a Type I error. The chance of at least one Type I error then depends on how many null hypotheses are true and on the dependence among their test decisions.

A sound plan starts before looking at results. Researchers can identify the main question and outcomes in advance, specify the tests they intend to conduct, and report how many tests were performed. If they explore many outcomes and then highlight only a small p-value, readers need to know about that broader search to understand the risk of false positives.

Common Mistakes and AP Exam Tip

  • Calling the overall probability “0.05”: The per-test \(\alpha\) is \(0.05\), but in the 20-test independent all-null example, the chance of at least one Type I error is about \(0.6415\).
  • Adding the probabilities and reporting \(20(0.05)=1\) as a probability: One is the expected number of Type I errors in 20 tests, not the probability of at least one. A probability cannot exceed 1. Use the complement calculation \(1-(1-\alpha)^m\) for the independent all-null case.
  • Forgetting the assumptions: The formula \(1-(1-\alpha)^m\) requires independent tests with the same \(\alpha\), and it describes at least one Type I error when all the null hypotheses are true. State these conditions when interpreting the result.
  • Saying every significant result must be false: Multiple testing increases the chance of at least one false positive; it does not establish that any particular significant result is false.
  • Changing alpha after seeing the results: Decisions about the number of tests and significance levels should be made as part of a plan, not adjusted to make observed results appear more convincing.

For full-credit communication, name the individual \(\alpha\), state how many tests are included, identify the assumptions, show the complement calculation, and interpret the result as a chance of at least one Type I error across the collection. Do not describe it as the probability that a particular null hypothesis is true.

Key takeaway: When many independent tests are run at the same significance level and all null hypotheses are true, the chance of at least one Type I error is \(1-(1-\alpha)^m\). With 20 tests at \(\alpha=0.05\), that chance is about \(0.6415\), not \(0.05\).

Check Your Understanding

Answer each question using the assumptions stated in the question.

  1. In your own words, what is the familywise Type I error probability?
  2. For 8 independent tests with every null hypothesis true and \(\alpha=0.05\), write the expression for the probability of at least one Type I error. You do not need to evaluate it.
  3. In the 20-test example, what does \(20(0.05)=1\) describe, and why is it not the probability of at least one Type I error?
  4. Why might \(1-(1-\alpha)^m\) not give the exact overall probability when test decisions are dependent?
  5. What per-test significance level did the final worked example calculate for 20 independent tests and a \(0.05\) familywise target, assuming all null hypotheses are true?