Tutorials › AP Statistics › How Significance Level Affects Power and Type II Error

Inference decisions and errors · Tutorial 591 of 1000

How Significance Level Affects Power and Type II Error

Compare alpha, beta, and power for the same one-proportion test to see why a more demanding rejection rule usually lowers power.

Intermediate 9 min read

What You'll Learn

  • Explain how significance level determines a test’s rejection cutoff.
  • Estimate beta and power at a specified alternative proportion.
  • Compare power when alpha changes and other test details stay fixed.
  • Describe the tradeoff between Type I and Type II errors.
  • Interpret a power estimate in context and identify its approximation.

Why Changing Alpha Affects Power

In “How Sample Size Affects Power,” the significance level stayed fixed while the sample size changed. Here we hold the sample size, hypotheses, and specified alternative fixed, and change only the significance level, \(\alpha\). This reveals a different part of the power story: the test’s rejection rule becomes more or less demanding as alpha changes.

We will consider an upper-tail one-proportion z test of whether a text-message reminder increases the proportion of eligible residents who complete an online community-service registration. The null hypothesis is \(H_0:p=0.25\), and the alternative is \(H_a:p>0.25\). We will estimate power at the specified alternative \(p_1=0.32\), using \(n=200\) residents.

Key idea: With the hypotheses, sample size, and specified alternative held fixed, raising \(\alpha\) makes it easier to reject \(H_0\). The test then has greater power at that specified alternative and a smaller \(\beta\). Lowering \(\alpha\) makes rejection harder, usually lowering power and increasing \(\beta\).

Recall from the earlier tutorials on significance level, beta, and power that \(\alpha\) is the probability of a Type I error, \(\beta(p_1)\) is the probability of failing to reject \(H_0\) when the specified alternative \(p_1\) is true, and power is \(1-\beta(p_1)\). These probabilities refer to repeated use of the test under the stated truth.

For an upper-tail test, the approximate rejection cutoff for \(\hat{p}\) is set using the null model. If the test uses significance level \(\alpha\), its critical value \(z^*\) is chosen so that the area to its right under the standard normal curve is \(\alpha\). To estimate power, find the probability of reaching the cutoff under the sampling distribution centered at the specified alternative \(p_1\).

$$ \text{Cutoff}=p_0+z^*\sqrt{\frac{p_0(1-p_0)}{n}} \qquad \text{Power}(p_1)\approx P(\hat{p}\geq \text{Cutoff}\mid p=p_1) $$

A larger \(\alpha\) gives a smaller upper-tail critical value \(z^*\), so the cutoff moves closer to \(p_0\). More sample proportions then fall in the rejection region. That raises the probability of rejecting \(H_0\) when \(p_1\) is true, which means higher power and lower \(\beta\). This does not mean that lowering alpha is wrong: it reflects a deliberate choice to make a Type I error less likely, at the cost of making a Type II error more likely for the specified alternative.

A Numerical Comparison

Worked Example: Comparing Alpha of 0.05 and 0.10

A community program will test whether its reminder increases the proportion of eligible residents who complete registration. The test uses \(H_0:p=0.25\), \(H_a:p>0.25\), and \(n=200\). Estimate and compare power at \(p_1=0.32\) when \(\alpha=0.05\) and when \(\alpha=0.10\).

State: We are comparing the probability of rejecting \(H_0:p=0.25\) if the true registration-completion proportion is \(0.32\). The two plans differ only in alpha. Their sample size, hypotheses, and specified alternative are the same.

Plan and conditions: Use the null distribution to find the cutoff for each alpha, then use the sampling distribution at \(p_1=0.32\) to estimate power. Assume the 200 residents are selected through an appropriate random process. If sampling is without replacement, suppose the eligible population contains at least 2,000 residents; then \(200\leq0.10(2000)\), so the 10% condition is met. Under the null, the Large Counts checks are \(np_0=200(0.25)=50\) and \(n(1-p_0)=200(0.75)=150\), both at least 10. Under the specified alternative, \(np_1=200(0.32)=64\) and \(n(1-p_1)=200(0.68)=136\), also both at least 10. These conditions support the normal approximations for this estimate.

Do: The null standard deviation of \(\hat{p}\) is \(\sqrt{0.25(0.75)/200}=\sqrt{0.0009375}\approx0.0306186\). For \(\alpha=0.05\), use \(z^*=1.6449\), rounded to four decimal places. The cutoff is \(0.25+1.6449(0.0306186)\approx0.30036\).

Under \(p_1=0.32\), the standard deviation is \(\sqrt{0.32(0.68)/200}=\sqrt{0.001088}\approx0.0329848\). Standardizing the cutoff using this alternative distribution gives \(z=(0.30036-0.32)/0.0329848\approx-0.5954\). Thus the estimated power is \(P(Z\geq-0.5954)\approx0.7242\), and the estimated Type II error probability is \(\beta\approx1-0.7242=0.2758\). These probabilities are rounded to four decimal places.

For \(\alpha=0.10\), use \(z^*=1.2816\). The null-based cutoff is \(0.25+1.2816(0.0306186)\approx0.28924\). Using the same alternative standard deviation, the standardized cutoff is \(z=(0.28924-0.32)/0.0329848\approx-0.9326\). Therefore, estimated power is \(P(Z\geq-0.9326)\approx0.8245\), and estimated \(\beta\) is \(1-0.8245=0.1755\).

Conclude: If the true proportion is \(0.32\), the test has estimated power of about \(0.7242\) at \(\alpha=0.05\) and \(0.8245\) at \(\alpha=0.10\). Raising alpha increases estimated power by about \(0.1003\) and lowers estimated \(\beta\) by the same amount. The more permissive rejection rule is more likely to detect this specified increase, but it also allows a greater probability of a Type I error.

What the Cutoffs Show

The two cutoffs make the change in the rejection rule visible. At \(\alpha=0.05\), the approximate test rejects for sample proportions of about \(0.30036\) or greater. At \(\alpha=0.10\), it rejects for sample proportions of about \(0.28924\) or greater. The second cutoff is closer to the null proportion, \(0.25\), so more possible sample results lead to rejection.

Under the specified truth \(p_1=0.32\), a sample proportion is more likely to exceed the lower cutoff. Equivalently, less of the alternative sampling distribution falls below that cutoff, so \(\beta\), the probability of failing to reject a false null at \(p_1=0.32\), decreases. Power and \(\beta\) are complements for the same specified alternative: if one rises, the other falls.

Alpha itself is not the probability that the null hypothesis is true, nor is it the probability that a particular rejection is an error. It is the probability of rejecting \(H_0\) when \(H_0\) is true, for a test procedure with that chosen significance level. The comparison concerns the test’s long-run behavior under different possible truths; it does not tell us which hypothesis is true for one particular sample.

Worked Example: Using a More Stringent Alpha of 0.01

Keep the same registration study, hypotheses, sample size, and specified alternative, but set \(\alpha=0.01\). Estimate power and \(\beta\), and compare them with the results for \(\alpha=0.05\).

State: We want the probability of rejecting \(H_0:p=0.25\) if the true proportion is \(p_1=0.32\), using a more stringent significance level of \(0.01\).

Plan and conditions: Use an upper-tail normal cutoff based on \(p_0=0.25\), then calculate the probability above that cutoff under \(p_1=0.32\). The random-process assumption and, if applicable, the 10% condition are as in the first example. The null Large Counts values remain \(50\) and \(150\); the alternative values remain \(64\) and \(136\). All are at least 10, so the normal approximation is supported.

Do: For \(\alpha=0.01\), \(z^*=2.3263\). Using the null standard deviation \(0.0306186\), the cutoff is \(0.25+2.3263(0.0306186)\approx0.32123\). Under \(p_1=0.32\), the standard deviation is \(0.0329848\), as before. The standardized cutoff is \(z=(0.32123-0.32)/0.0329848\approx0.0373\). Estimated power is \(P(Z\geq0.0373)\approx0.4851\), so estimated \(\beta\) is \(1-0.4851=0.5149\), rounded to four decimal places.

Conclude: If the true completion proportion is \(0.32\), the test with \(\alpha=0.01\) has estimated power of about \(0.4851\). That is lower than the estimated power of \(0.7242\) at \(\alpha=0.05\), while estimated \(\beta\) rises from \(0.2758\) to \(0.5149\). The stricter rule reduces the chance of a Type I error, but in this setting it also makes the specified increase harder to detect.

Keeping the Comparison Fair

A useful comparison changes alpha while holding the other ingredients fixed: the null and alternative hypotheses, sample size, and specified true value \(p_1\). If those ingredients change too, the resulting power difference cannot be attributed to alpha alone. For example, changing \(p_1\) changes how far the truth is from the null, which also affects power.

The estimates above use normal approximations. A sample proportion is discrete, so an actual one-proportion z test’s Type I error probability and power can differ slightly from these continuous-model estimates. The direction of the comparison is the important point: for this same upper-tail test and specified alternative, increasing alpha lowers the cutoff and raises estimated power. A calculator can be used to evaluate the normal tail probabilities, but the alternative distribution—not the null distribution—is the one used for the power calculation.

Key takeaway: At a fixed sample size and specified alternative, a larger significance level generally increases power and decreases \(\beta\), while increasing the probability of a Type I error when \(H_0\) is true. A smaller alpha reverses that tradeoff. Always state the specified alternative when interpreting power or beta.

Common Mistakes and AP Exam Tip

  • Thinking alpha is the probability that \(H_0\) is true: Alpha is the probability of rejecting a true \(H_0\), not a probability assigned to the truth of the hypothesis.
  • Using the null distribution to calculate power: The null model sets the rejection cutoff. To estimate power at \(p_1\), calculate the probability of crossing that cutoff using the sampling distribution centered at \(p_1\).
  • Forgetting that beta depends on the alternative: A value of \(\beta\) must be tied to a specified true proportion. A complete interpretation names that value and the test’s failure-to-reject probability.
  • Claiming that higher alpha improves every part of a test: Higher alpha raises power here, but also raises the risk of a Type I error. It is a tradeoff, not a cost-free improvement.
  • Changing several features at once: To isolate alpha’s effect, keep the sample size, hypotheses, and \(p_1\) fixed and clearly identify the alpha values being compared.
  • Calling a normal approximation exact: The sample proportion is discrete. Describe these power and beta values as estimates when they come from the normal approximation.

For a full-credit comparison, say what was held fixed, identify the specified alternative, report how power and beta changed, and interpret the probabilities in context. For instance: “If the true registration-completion proportion is \(0.32\), raising \(\alpha\) from \(0.05\) to \(0.10\) raises estimated power from \(0.7242\) to \(0.8245\), while raising the nominal significance level from \(0.05\) to \(0.10\); the actual Type I error probability for this discrete test may differ from those levels.”

Check Your Understanding

Use the relationship among alpha, beta, and power to answer each question.

  1. Why must the hypotheses, sample size, and specified alternative stay fixed to isolate alpha’s effect on power?
  2. For an upper-tail test, how does increasing alpha affect the critical value and the approximate rejection cutoff?
  3. In the registration example, what does estimated power of \(0.8245\) at \(p_1=0.32\) mean in context?
  4. At the same specified alternative, why does a decrease in estimated beta correspond to an increase in estimated power?
  5. Why might an actual one-proportion z test’s Type I error probability or power differ slightly from a normal-approximation estimate?