Tutorials › AP Statistics › Power and Type II Error Probability

Inference errors and practical significance · Tutorial 792 of 1000

Power and Type II Error Probability

Use the complement of beta to find and explain a test’s power for a specified false-null value.

Intermediate 8 min read

What You'll Learn

  • Explain why power equals one minus beta for the same specified true mean and test.
  • Calculate power from a supplied Type II error probability.
  • Interpret beta and power as complementary repeated-sample outcomes in context.
  • Convert simulation counts of failures to reject into beta and power.
  • Avoid confusing beta with alpha or with the probability that a hypothesis is true.

Power and Type II Error Probability Are Complements

In “What Power of a Test Means,” you learned that power is the probability a test rejects \(H_0\) when a particular false-null value is true. The matching probability for failing to reject \(H_0\) under that same false-null value is called beta, written \(\beta\). Together, these two probabilities describe the two possible decisions the test can make when \(H_0\) is false.

Definition: For a specified true value of the population mean that makes \(H_0\) false, \(\beta\) is the probability that the test fails to reject \(H_0\). Power is the probability that the test rejects \(H_0\) for that same true value and test design.

When that particular false-null value is true, the test either rejects \(H_0\) or fails to reject \(H_0\). These outcomes are mutually exclusive and cover all possible test decisions. Therefore, their probabilities add to 1. Since power is the probability of rejection and \(\beta\) is the probability of failing to reject, they are complements.

$$ \text{Power}=1-\beta \qquad\text{and}\qquad \beta=1-\text{Power}. $$

For example, if \(\beta=0.24\) for a specified true mean, power is \(1-0.24=0.76\). If that mean really is the population mean, the test would fail to reject \(H_0\) in about 24% of repeated samples and reject \(H_0\) in about 76% of repeated samples. The two percentages add to 100%.

The phrase for the same specified true value and test design matters. As discussed in “What Power of a Test Means,” power depends on the population mean assumed to be true and on the test being used. The corresponding \(\beta\) depends on those same details. You cannot take a \(\beta\) calculated for one alternative mean and use it to describe power at another.

Formula: If the supplied Type II error probability is \(\beta\), calculate power by subtracting it from 1. If the supplied probabilities are percentages, first express \(\beta\) as a proportion, or subtract its percentage directly from 100%.

Interpreting the Two Complementary Outcomes

A Type II error is failing to reject a false null hypothesis. Thus, \(\beta\) describes the chance of that particular error when a specified false-null value is true. Power describes the chance of the other decision—rejecting \(H_0\)—under that same population truth. A large \(\beta\) means the test often misses that specified departure from the null; a large power means the test often detects it by rejecting \(H_0\).

These interpretations are about repeated use of a testing procedure, not certainty about one future sample. If power is 0.76, it does not mean that 76% of any particular sample will lead to rejection. It means that, over repeated samples under the stated model and design, about 76% of the tests would reject \(H_0\), assuming the specified mean is truly the population mean.

Beta and power are not probabilities that hypotheses are true or false. A value such as \(\beta=0.24\) does not mean there is a 24% chance that \(H_0\) is true, and power of 0.76 does not mean there is a 76% chance that the alternative is true. The probabilities are conditional on a particular population truth; they describe how the procedure would behave under that condition.

Key distinction: For a specified false-null mean, \(\beta\) is the probability of failing to reject \(H_0\), while power is the probability of rejecting \(H_0\). They are complementary probabilities, not competing estimates of whether a hypothesis is true.

Beta is also not the same as alpha. As explained in “Alpha as the Type I Error Rate in t Tests,” \(\alpha\) is the probability of rejecting \(H_0\) when \(H_0\) is true. In contrast, \(\beta\) concerns failing to reject \(H_0\) when a specified false-null value is true. The two probabilities refer to different population situations and different kinds of error.

Worked Examples: Finding Power from Beta

Worked Example: A Mean Battery-Life Test

A fictional manufacturer plans a two-sided one-sample t test of \(H_0:\mu=12\) hours against \(H_a:\mu\ne12\), using a specified sample size and \(\alpha=0.05\). Here, \(\mu\) is the true mean battery life for the population represented by the sampling process. Planning calculations give \(\beta=0.18\) when the true mean battery life is 13.5 hours. Find and interpret the power for that specified true mean.

State. The value 13.5 hours is the false-null mean being considered. The supplied \(\beta\) is the probability of failing to reject \(H_0\) if this is the true population mean.

Plan. For this same test and true mean, power is the complementary probability: \(1-\beta\). No new test statistic or p-value is needed because beta has already been supplied.

Do. Substitute the given value of beta:

$$ \text{Power}=1-\beta=1-0.18=0.82. $$

Conclude. If the true mean battery life is 13.5 hours, this test would reject the claim that the population mean is 12 hours in about 82% of repeated samples, assuming the stated test design and conditions apply. In about 18% of those repeated samples, the test would fail to reject \(H_0\), even though the true mean is 13.5 hours.

Worked Example: Beta Given as a Percentage

A fictional school nutrition team plans a one-sided one-sample t test of \(H_0:\mu=20\) against \(H_a:\mu<20\), where \(\mu\) is the mean number of minutes students in the target population spend eating lunch. For a particular planned design, the Type II error probability is 32% when the true mean is 18 minutes. What is the power, and how should it be interpreted?

Identify the matching situation. The 32% beta refers to this test when the true mean is 18 minutes, not to every possible mean below 20. The power sought must refer to that same test and true mean.

Calculate the complement. The supplied beta as a proportion is \(0.32\). Therefore:

$$ \text{Power}=1-0.32=0.68. $$

Equivalently, subtract the percentage from 100%: \(100\%-32\%=68\%\). If the true mean lunch time is 18 minutes, the test would reject \(H_0:\mu=20\) in about 68% of repeated samples. It would fail to reject \(H_0\) in about 32% of repeated samples. The direction of the alternative is consistent with the specified mean being below 20 minutes.

The result describes the planned procedure’s long-run behavior under the assumption that the mean is 18 minutes. It is not a claim that the mean actually is 18 minutes, nor that the test has 68% power for every mean less than 20 minutes.

Worked Example: Turning Simulation Results into Beta and Power

A fictional environmental group simulates a planned two-sided t test of a mean water-quality measurement. The null hypothesis specifies a mean of 40 units. In 500 simulated studies where the true mean is 43 units, the test fails to reject \(H_0\) in 145 studies. Estimate beta and power for a true mean of 43 units.

Find the estimated beta. The simulated Type II errors are the failures to reject \(H_0\) when the specified false-null mean is true. The estimated proportion is:

$$ \widehat{\beta}=\frac{145}{500}=0.29. $$

Find the estimated power. By complementation:

$$ \widehat{\text{Power}}=1-0.29=0.71. $$

As a check, \(500-145=355\) of the simulated studies rejected \(H_0\), and \(355/500=0.71\). Both calculations give the same estimated power.

Interpret in context. If the true mean water-quality measurement is 43 units, the planned test would reject \(H_0\) in about 71% of repeated studies represented by the simulation. It would fail to reject the null claim in about 29%. These are simulation-based estimates from 500 repetitions, rounded to two decimal places; they are not guarantees about a future study.

Worked Example: Beta Changes with the Specified True Mean

A fictional sports-science team plans the same one-sided test of \(H_0:\mu=30\) against \(H_a:\mu>30\) for mean sprint-recovery time in seconds, using a specified design. Planning calculations give \(\beta=0.25\) when the true mean is 32 seconds and \(\beta=0.08\) when the true mean is 35 seconds. Find and compare power at these two specified means.

At a true mean of 32 seconds:

$$ \text{Power}=1-0.25=0.75. $$

If the true mean recovery time is 32 seconds, the test would reject \(H_0\) in about 75% of repeated samples and fail to reject it in about 25%.

At a true mean of 35 seconds:

$$ \text{Power}=1-0.08=0.92. $$

If the true mean recovery time is 35 seconds, the test would reject \(H_0\) in about 92% of repeated samples and fail to reject it in about 8%.

The two power values differ because each is tied to a different specified population mean. In this example, the mean of 35 seconds is farther above the null value of 30 seconds than is 32 seconds, and the supplied beta is smaller at 35. Thus, the test is more likely to reject the null at that specified value. Do not report a single power value without saying which true mean it describes.

Common Mistakes and AP Exam Tips

  • Subtracting beta from alpha: The formula is \(\text{Power}=1-\beta\), not \(\alpha-\beta\). Alpha describes Type I error under a true null; beta describes Type II error at a specified false-null value.
  • Calling beta the probability of a Type I error: Beta is the probability of failing to reject a false null hypothesis. Rejecting a true null is a Type I error and is described by alpha.
  • Leaving out the specified true mean: A complete interpretation says which false-null value is assumed true. Power and beta can differ at different population means for the same test.
  • Interpreting power as the probability that the alternative is true: Power is conditional on the specified alternative value being true. It describes repeated-sample decisions, not the probability of a hypothesis.
  • Forgetting that beta and power refer to opposite decisions: Under the same specified false-null truth, beta counts failures to reject and power counts rejections. State the corresponding decision when interpreting each one.
  • Treating a simulated proportion as exact: If beta is estimated from simulated studies, then the complement estimates power. Identify it as an estimate and keep the denominator clear.

For a full-credit interpretation, connect the probability to the test decision, the specified population mean, and the context. One useful structure is: “If the true population mean is [value], this test would reject \(H_0\) in about [power] of repeated samples; it would fail to reject \(H_0\) in about [beta] of those samples, assuming the stated design and conditions hold.”

Key takeaway: For a specified true mean that makes \(H_0\) false, power and beta describe complementary test outcomes: \(\text{Power}=1-\beta\). Interpret both as long-run probabilities for that particular population truth and test design.

Check Your Understanding

For each question, keep the specified true population value and test decision in view.

  1. A test about a population mean has \(\beta=0.14\) when the true mean is 52 units. What is the power, and how would you interpret it?
  2. A planning study reports \(\beta=27\%\) for a specified false-null mean. State the power as a percentage and explain what the two percentages count.
  3. In 400 simulated studies under a specified false-null mean, the test fails to reject \(H_0\) in 96 studies. Find the estimated beta and estimated power.
  4. Why is it incomplete to report “power is 0.80” without stating the population mean assumed to be true?
  5. Explain why beta is not the same as alpha, even though both are probabilities related to errors in testing.