Tutorials › AP Statistics › Multiple Testing and False Positives

Inference errors and practical significance · Tutorial 797 of 1000

Multiple Testing and False Positives

Learn how the number of tests affects false-positive risk, and why a 0.05 significance level for each test does not mean a 0.05 risk across the whole set.

Intermediate 9 min read

What You'll Learn

  • Explain what a false positive means when testing a population mean.
  • Distinguish the Type I error rate for one test from the chance of at least one false positive across many tests.
  • Calculate the expected number of false positives when several true null hypotheses are tested at alpha = 0.05.
  • Calculate the chance of one or more false positives when the tests are independent and all null hypotheses are true.
  • Explain why dependence among tests affects the probability calculation.
  • Describe why an isolated significant result among many tests needs cautious interpretation.

Why the Number of Tests Matters

In “Type I Error in a Test About Means” and “Alpha as the Type I Error Rate in t Tests,” you learned that a Type I error occurs when a test rejects a true null hypothesis. At a significance level of \(\alpha=0.05\), a test has a 5% long-run chance of a Type I error when its null hypothesis is true. That statement describes one test.

Researchers may test several population means in one study: for example, whether a program affects several different outcomes, or whether mean results differ across many groups. If every null hypothesis is actually true, each test still has its own 0.05 Type I error rate. But the chance that at least one test in the collection produces a false positive can be much larger than 0.05.

Definition: A false positive is a test result that rejects a true null hypothesis. When several hypotheses are tested as a group, the family-wise error rate is the probability of making at least one Type I error in that group.

“Family” here means the set of tests considered together. The family-wise error rate is not the same as the Type I error rate for an individual test. This distinction is the central idea: keeping each test at \(\alpha=0.05\) does not keep the chance of any false positive across many tests at 0.05.

Calculating the Chance of at Least One False Positive

Suppose all \(m\) null hypotheses are true, every test uses \(\alpha=0.05\), and the tests are independent. The probability that a single test does not produce a false positive is \(1-0.05=0.95\). Independence lets us multiply these probabilities to find the chance that none of the \(m\) tests produces a false positive. Subtract that result from 1 to get the chance of at least one.

$$ P(\text{at least one false positive}) =1-P(\text{no false positives}) =1-(1-\alpha)^m $$

At \(\alpha=0.05\), this becomes \(1-(0.95)^m\). As the number of independent tests grows, the chance of at least one false positive grows too. Each test still has a 0.05 Type I error rate; it is the chance across the set that has increased.

Important condition: The formula \(1-(1-\alpha)^m\) for the chance of at least one false positive assumes that all \(m\) null hypotheses are true and the tests are independent. If the tests are related, this multiplication may not give the correct probability.

The tests might not be independent if they use the same participants or measure outcomes that are closely related. For example, measurements of several similar skills from the same students may tend to move together. In that case, the probability of at least one false positive depends on how the test results are related, not just on \(m\) and \(\alpha\).

There is a useful calculation that does not require independence: if all \(m\) null hypotheses are true and each test has Type I error probability 0.05, the expected number of false positives is \(0.05m\). To see why, count each false positive as 1 and each non-false positive as 0. Each test contributes an expected \(0.05\) to the count, so the expected total is the sum of those contributions. This does not tell us the probability of at least one false positive, and it does not mean we will observe exactly that expected number.

$$ \text{Expected number of false positives}=m\alpha $$

Worked Example: Twelve Tests of Mean Outcomes

Worked Example: A School District’s New Study Routine

Suppose a fictional school district tests whether a new study routine changes the population mean score on each of 12 different academic measures. For this illustration, assume all 12 null hypotheses are true, each test uses \(\alpha=0.05\), and the test results are independent. What is the expected number of false positives, and what is the probability of at least one?

State. Find both the expected count of Type I errors among the 12 tests and the chance that one or more tests reject a true null hypothesis.

Plan. Use \(m\alpha\) for the expected count. Use \(1-(1-\alpha)^m\) for the probability of at least one false positive; this second calculation is appropriate because independence and the truth of all 12 null hypotheses are assumed.

Do. The expected number is \(12(0.05)=0.60\) false positives. The chance of no false positives is \(0.95^{12}=0.5404\), rounded to four decimal places. Therefore, the chance of at least one is \(1-0.95^{12}=1-0.5404=0.4596\), rounded to four decimal places.

Conclude. Under these assumptions, the district should expect 0.60 false positives on average across repeated sets of 12 tests. The probability of at least one false positive in a single set is about 0.4596, or 46.0%. This is much higher than the 0.05 Type I error rate for any one test. An expected count of 0.60 does not mean a fraction of a false positive occurs in a particular study.

Not Every Test in a Study Is a True Null

The calculation needs careful interpretation when some null hypotheses are false. A Type I error can occur only for a true null hypothesis. If a null hypothesis is false, rejecting it is not a false positive; failing to reject it could instead be a Type II error. For the expected number of false positives, count the true null hypotheses, rather than automatically counting every test.

For the probability of at least one false positive, the independence formula applies to the set of true null hypotheses if those tests are independent. It does not require the false null hypotheses to be included in that calculation. In practice, researchers may not know which null hypotheses are true, so these calculations are often illustrations of what can happen under stated assumptions, rather than a way to label a particular result as false.

Worked Example: Five of Eight Mean Tests Have True Nulls

A fictional sports-science team tests eight mean outcomes related to a training plan. Assume five of the null hypotheses are true, three are false, and all eight tests are independent. Each test uses \(\alpha=0.05\). Find the expected number of false positives and the chance of at least one false positive.

State. Calculate the false-positive risk for the five tests whose null hypotheses are true.

Plan. False positives are rejections of true null hypotheses, so use \(m=5\), not 8, in these calculations. Use \(m\alpha\) for the expected count and \(1-(1-\alpha)^m\) for the probability of at least one, given independence.

Do. The expected number is \(5(0.05)=0.25\). The probability that none of the five true nulls is rejected is \(0.95^5=0.7738\), rounded to four decimal places. So the probability of at least one false positive is \(1-0.95^5=1-0.7738=0.2262\), rounded to four decimal places.

Conclude. Under these assumptions, the expected number of false positives is 0.25, and the chance of at least one false positive is about 0.2262, or 22.6%. The three false null hypotheses do not count toward the false-positive risk: rejecting a false null is not a Type I error.

Worked Example: Twenty Independent Tests

Worked Example: Comparing Several Environmental Measures

A fictional environmental team tests 20 different population means, such as average readings for different water-quality measures. Suppose all 20 null hypotheses are true, all tests are independent, and each test uses \(\alpha=0.05\). Find the expected number of false positives and the probability of at least one.

State. Assess the chance of false positives across the entire set of 20 tests under the assumption that every null hypothesis is true.

Plan. Apply \(m\alpha\) to find the expected count, and \(1-(1-\alpha)^m\) to find the chance of at least one. The independence assumption is required for the second formula.

Do. The expected number is \(20(0.05)=1.00\). For the probability of at least one, \(0.95^{20}=0.3585\), rounded to four decimal places, so \(1-0.95^{20}=1-0.3585=0.6415\), rounded to four decimal places.

Conclude. Under these assumptions, the expected number of false positives is 1.00, and the chance of at least one is about 0.6415, or 64.2%. The expected count does not guarantee exactly one false positive in this set. It describes the long-run average across repeated sets of 20 tests.

What a Significant Result Does—and Does Not—Show

If one test in a large set has a p-value below 0.05, that result is statistically significant at the 0.05 level for that test. It does not follow that the result is a false positive. The calculations above describe probabilities over repeated testing under assumptions, especially that a particular null hypothesis is true. They cannot determine whether an individual null hypothesis is true.

At the same time, a significant result selected from many tests deserves context. If a study tests many outcomes and highlights only the significant ones, readers may not see how many tests were conducted. An isolated small p-value may be less persuasive as a stand-alone finding when it emerged from a large set of tests than when it came from a single, planned test. Researchers can make the testing process easier to evaluate by explaining which outcomes were planned in advance and reporting the full set of tests, not only the significant results.

This is one reason to distinguish exploratory analysis from a test of a planned claim. Exploring many outcomes can help identify questions for future study, but an interesting result from that exploration can be checked in a new study. A new study does not make the original result automatically true or false; it provides additional evidence. Always describe the result in the context of the design and the population parameter, as emphasized in “Tying the Conclusion to the Original Claim.”

Common Mistakes and AP Exam Tips

  • Confusing per-test alpha with family-wise error rate: An \(\alpha\) of 0.05 gives a 0.05 Type I error probability for one test when its null is true. It does not say that the chance of at least one false positive across many tests is 0.05.
  • Calling the expected number a probability: For 12 true null hypotheses, \(12(0.05)=0.60\) is an expected count, not a 60% chance of at least one false positive. Under independence, that chance is \(1-0.95^{12}=0.4596\).
  • Using the independence formula without checking its assumption: Multiplying the probabilities of no false positives requires independent tests. If outcomes or test results are related, do not claim the formula gives the exact family-wise error rate.
  • Counting false nulls as possible Type I errors: A Type I error is rejecting a true null. If only five of eight nulls are true, the expected number of false positives at \(\alpha=0.05\) is \(5(0.05)=0.25\), not \(8(0.05)\).
  • Claiming a particular significant result must be false: A higher chance of at least one false positive across many tests does not identify which result, if any, is false. A full-credit answer describes the long-run chance under the stated assumptions and does not label a specific result without evidence.
  • Forgetting to state the assumptions: When using \(1-(1-\alpha)^m\), say that the tests are independent and the \(m\) null hypotheses being counted are true. Without those assumptions, the calculation may not apply.

A strong explanation names the number of tests, identifies which null hypotheses are assumed true, states whether independence is assumed, and distinguishes the expected count from the probability of at least one error. For example: “If all 12 null hypotheses are true and the tests are independent, the probability of at least one false positive is \(1-0.95^{12}=0.4596\).”

Key takeaway: At \(\alpha=0.05\), each true null hypothesis has a 5% Type I error rate, but testing many hypotheses can make at least one false positive much more likely. Under independence, the probability across \(m\) true nulls is \(1-(0.95)^m\); the expected number is \(0.05m\).

Check Your Understanding

Use \(\alpha=0.05\) unless a question says otherwise. State any assumptions needed for a probability calculation.

  1. In one sentence, distinguish a test’s Type I error rate from the family-wise error rate for a set of tests.
  2. For 6 independent tests with all null hypotheses true, what is the expected number of false positives? Show the calculation.
  3. For those 6 tests, write the expression for the probability of at least one false positive. What assumption makes this formula appropriate?
  4. In a set of 10 tests, suppose 4 null hypotheses are true and the tests are independent. Which number of tests should be used to calculate the expected number of false positives, and why?
  5. A study reports one significant result after testing many outcomes. Why does that not prove the result is a false positive, and what context would help readers interpret it?