Tutorials › AP Statistics › How Significance Level Affects Power and Errors

Inference errors and practical significance · Tutorial 794 of 1000

How Significance Level Affects Power and Errors

Learn why a higher significance level makes a t test more likely to reject, increasing its Type I error rate and power while reducing Type II error probability for a specified true mean.

Intermediate 10 min read

What You'll Learn

  • Explain how increasing alpha expands a t test’s rejection region.
  • Describe the long-run Type I error tradeoff when alpha changes.
  • Compare power and Type II error probability at different alpha levels for a specified true mean.
  • Distinguish a changed decision on one set of data from changed long-run test behavior.
  • Interpret power and error rates in context without treating them as guarantees.

The Tradeoff When Alpha Changes

In “How Sample Size Affects Power in a t Test,” the significance level was held fixed while sample size changed. Here we hold the study design and assumed true mean fixed and ask what happens when the significance level, \(\alpha\), changes. The key is that \(\alpha\) sets how much evidence against \(H_0\) is required before a test rejects it.

For a t test, a larger \(\alpha\) makes the rejection region larger. For a two-sided test, the test rejects for less extreme values in either tail. For a one-sided test, the cutoff in the direction of the alternative moves closer to the null value. Either way, when \(H_0\) is true, a larger rejection region makes a Type I error more likely in repeated testing.

Key idea: With the test design and direction held fixed, increasing \(\alpha\) increases the long-run Type I error rate. For a specified false-null mean, it also increases power and decreases the Type II error probability, \(\beta\). These are long-run probabilities, not guarantees about one particular sample.

As covered in “Alpha as the Type I Error Rate in t Tests,” \(\alpha\) is the test’s long-run Type I error rate when the null hypothesis is true. In a t test that meets its conditions, the probability of rejection under \(H_0\) is controlled at the chosen significance level. Thus, raising \(\alpha\) from 0.01 to 0.05 accepts a higher long-run chance of rejecting a true null hypothesis.

For a specified true mean that makes \(H_0\) false, power is the probability of rejecting \(H_0\), and \(\beta\) is the probability of failing to reject it. As explained in “Power and Type II Error Probability,” these probabilities are complements: \(\text{Power}=1-\beta\). If the rejection region expands, more results fall into it under that false-null mean. Power therefore increases and \(\beta\) decreases.

$$ \text{Increasing }\alpha \quad\Longrightarrow\quad \text{larger rejection region} \quad\Longrightarrow\quad \begin{cases} \text{higher Type I error rate when }H_0\text{ is true}\\ \text{higher power and lower }\beta\text{ for a specified false }H_0 \end{cases} $$

This comparison only isolates the effect of \(\alpha\) if the other features remain fixed: the population parameter, test direction, sample size, variability, and true mean used to assess power. A power value is specific to that assumed true mean; a test does not have one power value that applies to every possible alternative.

How the Rejection Region Changes

Consider an upper-tailed one-sample t test of \(H_0:\mu=\mu_0\) versus \(H_a:\mu>\mu_0\). The test rejects when its t statistic exceeds a positive critical value. Raising \(\alpha\) lowers that critical value, so the sample mean does not have to be as far above \(\mu_0\) to enter the rejection region.

For a two-sided test, some of the allowed Type I error probability is assigned to each tail. Raising \(\alpha\) moves both critical values toward the null value. Results that were not extreme enough to reject at the smaller \(\alpha\) may now be in the rejection region. In both test types, this is why a higher \(\alpha\) increases the chance of rejecting \(H_0\).

Important distinction: Raising \(\alpha\) changes the rule for making a decision; it does not change the data, the p-value for those data, or the population mean. A p-value at or below the selected \(\alpha\) leads to rejection. The same p-value can therefore lead to different decisions at different significance levels.

In planning examples, the exact power of a t test depends on the distribution of its test statistic under the assumed alternative. A helpful AP-level approximation is to use a t critical value to locate the rejection boundary, then approximate the sample mean’s distribution under the specified true mean with a normal distribution. The examples below label these results as approximate. They use a planning estimate of the population standard deviation, rather than claiming it is known.

Worked Example: Raising Alpha in an Upper-Tailed t Test

Worked Example: Detecting a Mean Increase in Seedling Growth

A fictional greenhouse team plans a study of whether a lighting schedule increases the mean daily growth of seedlings. Let \(\mu\) be the population mean increase in growth, measured in millimeters. The team plans a one-sample t test of \(H_0:\mu=0\) versus \(H_a:\mu>0\), with \(n=100\). A pilot study gives a sample standard deviation of 10 millimeters, which the team uses as its planning estimate of the population standard deviation. Compare approximate power at \(\alpha=0.01\), \(0.05\), and \(0.10\), assuming the true mean increase is 2 millimeters.

State. The parameter is the population mean increase in daily seedling growth. The same test, sample size, planning variability, and assumed true mean will be used for all three significance levels. Only \(\alpha\) changes.

Plan. Use a one-sample t test. Assume the seedlings came from an appropriate random sample or randomized experiment, observations are independent, and the 10% condition is met if sampling without replacement. The data should not show severe skewness or outliers; with \(n=100\), the t procedure is generally robust to moderate departures from normality. For the power approximation, use the t critical value with \(df=99\) to find each rejection boundary. Then use a normal model for the sample mean centered at the assumed true mean of 2 millimeters.

Do. The planning standard error is:

$$ SE=\frac{10}{\sqrt{100}}=1\text{ millimeter}. $$

For an upper-tailed test, each critical value gives the boundary for the sample mean as \(0+t^*(1)\). With \(df=99\), the one-sided critical values for \(\alpha=0.01\), \(0.05\), and \(0.10\) are approximately 2.3646, 1.6604, and 1.2902. Under the planning approximation, \(\bar{x}\) has a normal distribution with mean 2 and standard deviation 1. The approximate power is the probability that \(\bar{x}\) exceeds the boundary:

  • At \(\alpha=0.01\), the boundary is \(0+2.3646(1)=2.3646\) millimeters. Thus, \(P(\bar{x}>2.3646)=P(Z>(2.3646-2)/1)=P(Z>0.3646)\approx0.3577\).
  • At \(\alpha=0.05\), the boundary is \(0+1.6604(1)=1.6604\) millimeters. Thus, \(P(\bar{x}>1.6604)=P(Z>(1.6604-2)/1)=P(Z>-0.3396)\approx0.6329\).
  • At \(\alpha=0.10\), the boundary is \(0+1.2902(1)=1.2902\) millimeters. Thus, \(P(\bar{x}>1.2902)=P(Z>(1.2902-2)/1)=P(Z>-0.7098)\approx0.7611\).

The approximate Type II error probabilities for this specified true mean are the complements of power: \(1-0.3577=0.6423\), \(1-0.6329=0.3671\), and \(1-0.7611=0.2389\), respectively.

Conclude. For a true mean increase of 2 millimeters, the approximate power rises from 0.3577 to 0.6329 to 0.7611 as \(\alpha\) increases from 0.01 to 0.05 to 0.10. The corresponding probability of failing to reject the false null hypothesis falls from 0.6423 to 0.3671 to 0.2389. In repeated studies under this specified alternative, the higher-alpha test would reject the null claim of no mean increase more often. The results are planning approximations, not predictions or guarantees for a single study.

Worked Example: One Data Set, Different Decisions

Worked Example: A p-Value Between Two Alpha Levels

A fictional city garden compares the mean height of seedlings grown under a new soil mixture with a benchmark height. A one-sample t test is conducted for \(H_0:\mu=18\) centimeters versus \(H_a:\mu>18\) centimeters. Suppose the test produces a p-value of 0.032. Consider the decision at \(\alpha=0.01\), \(0.05\), and \(0.10\).

Solution. The sample, test statistic, and p-value stay the same in all three comparisons. At \(\alpha=0.01\), \(0.032>0.01\), so fail to reject \(H_0\). At \(\alpha=0.05\), \(0.032\le0.05\), so reject \(H_0\). At \(\alpha=0.10\), \(0.032\le0.10\), so reject \(H_0\).

At the 0.05 and 0.10 levels, the data provide convincing evidence that the population mean seedling height under the new soil mixture is greater than 18 centimeters. At the 0.01 level, the data do not provide convincing evidence to reject the null claim using that stricter cutoff. This change in decision does not mean the population mean changed when \(\alpha\) changed; it means the decision rule changed.

A p-value is not the probability that \(H_0\) is true, and it is not itself the probability of a Type I error. The significance level is the long-run Type I error rate set for the test. A researcher who chooses a larger \(\alpha\) before seeing the data accepts a higher long-run risk of rejecting a true null hypothesis in exchange for greater power against specified alternatives.

Worked Example: What the Long-Run Rates Mean

Worked Example: Repeated Testing Under Two Population Truths

Return to the greenhouse test in the first example. Imagine, as a way to interpret long-run probabilities, 1,000 repeated studies with the same design. First suppose the null claim is true; then separately suppose the true mean increase is 2 millimeters. Compare \(\alpha=0.01\) with \(\alpha=0.10\).

Solution when the null is true. At \(\alpha=0.01\), the long-run Type I error rate is 0.01, so the expected number of rejections of a true null hypothesis in 1,000 studies is \(1000(0.01)=10\). At \(\alpha=0.10\), the expected number is \(1000(0.10)=100\). These are expected long-run counts, not a promise that exactly 10 or 100 errors will occur in any particular collection of studies.

Solution when the true mean increase is 2 millimeters. From the first example, approximate power is 0.3577 at \(\alpha=0.01\) and 0.7611 at \(\alpha=0.10\). Across 1,000 repeated studies under that assumed mean, the expected numbers of rejections are \(1000(0.3577)=357.7\), or about 358, and \(1000(0.7611)=761.1\), or about 761. The corresponding Type II error probabilities are 0.6423 and 0.2389, giving about \(1000(0.6423)=642\) and \(1000(0.2389)=239\) failures to reject the false null hypothesis.

Conclusion. The higher significance level increases the chance of rejecting \(H_0\) whether the null is true or false. When \(H_0\) is true, that increased chance is a higher Type I error rate. For the specified false-null mean, it is greater power and a lower Type II error probability. The repeated-study counts make the tradeoff visible, but the actual counts in a finite set of studies will vary.

Common Mistakes and AP Exam Tips

  • Claiming that a higher alpha only affects Type I error: It also changes power and \(\beta\) for a specified false-null mean. State all parts of the tradeoff: Type I error rate rises; power rises; Type II error probability falls.
  • Calling alpha the probability that this particular conclusion is wrong: \(\alpha\) describes a long-run Type I error rate when \(H_0\) is true. It is not a post-study probability that the null hypothesis is true or that a particular rejection is mistaken.
  • Forgetting to specify the alternative for power: Power must be interpreted for a particular true mean that makes \(H_0\) false. Name that mean and the response units, as in “for a true mean increase of 2 millimeters.”
  • Changing multiple features in a comparison: To attribute a power change to \(\alpha\), keep sample size, variability, true mean, and test direction fixed. Otherwise, more than one factor may explain the difference.
  • Confusing the p-value with alpha: The p-value is calculated from the sample and test; alpha is the cutoff selected for the decision. The p-value remains fixed when you compare it with different alpha levels.
  • Implying that a higher alpha guarantees rejection: A larger rejection region increases the long-run probability of rejection, but a particular sample can still fail to reject \(H_0\).
  • Describing an approximation as exact: When a planning standard deviation and normal approximation are used to estimate t-test power, identify the result as approximate and state the assumed true mean.

A strong AP response names the test and parameter, identifies what is being held fixed, and connects the larger rejection region to the error tradeoff. For a specified false-null mean, say that power increases and \(\beta\) decreases. When describing Type I error, state that its long-run rate increases when \(H_0\) is true. Keep each statement about repeated testing rather than promising an outcome for one study.

Key takeaway: Raising \(\alpha\) makes a t test more willing to reject \(H_0\). That increases the long-run Type I error rate when the null is true, but also increases power and reduces Type II error probability for a specified false-null mean.

Check Your Understanding

Assume the test design, sample size, variability, and specified true mean stay fixed unless a question says otherwise.

  1. A one-sided t test changes from \(\alpha=0.02\) to \(\alpha=0.05\). What happens to its rejection cutoff and long-run Type I error rate?
  2. For a specified true mean that makes \(H_0\) false, power is 0.72 at one significance level. What is \(\beta\) at that level?
  3. A test gives \(p=0.041\). State the reject-or-fail-to-reject decision at \(\alpha=0.01\) and at \(\alpha=0.05\). Does the p-value change?
  4. Why must a statement about power identify a particular true mean or mean difference?
  5. In repeated studies when \(H_0\) is true, what risk increases when a researcher raises \(\alpha\)? What benefit may the researcher gain when \(H_0\) is false?