Tutorials › AP Statistics › Common Misconceptions About Alpha, Beta, and Power

Inference decisions and errors · Tutorial 598 of 1000

Common Misconceptions About Alpha, Beta, and Power

Learn to interpret alpha, beta, and power correctly by stating what is being conditioned on and avoiding claims about the probability that a hypothesis is true.

Intermediate 9 min read

What You'll Learn

  • Explain why alpha is not the probability that the null hypothesis is true.
  • Interpret alpha as the probability of rejecting a true null hypothesis.
  • Interpret beta and power for a specified true alternative value.
  • Explain why “the probability of rejecting” is incomplete without stating what is true.
  • Distinguish alpha, power, and a p-value.
  • Describe repeated-sampling meanings of alpha and power without claiming they predict one individual test result.

Probability Depends on What We Assume Is True

The words “probability of rejecting” can sound straightforward, but a test’s rejection probability depends on the actual population condition. In “Statistical Conclusions and Possible Mistakes in Context,” we separated a test decision from the truth about the population. Here, we sharpen that distinction: alpha, beta, and power are conditional probabilities. Each describes the chance of a test decision given a specified truth.

For a one-proportion test with null hypothesis \(H_0:p=p_0\), alpha describes the probability of rejecting \(H_0\) when \(p=p_0\) is true. For a specified alternative value \(p_1\), beta describes the probability of failing to reject \(H_0\) when \(p=p_1\) is true. Power is the probability of rejecting \(H_0\) when that specified alternative value is true.

Definitions: For a test with \(H_0:p=p_0\), the significance level is \(\alpha=P(\text{reject }H_0\mid p=p_0)\). For a specified value \(p_1\) under the alternative, \(\beta(p_1)=P(\text{fail to reject }H_0\mid p=p_1)\), and the power at \(p_1\) is \(1-\beta(p_1)=P(\text{reject }H_0\mid p=p_1)\).

The vertical bar in probability notation means “given that.” Thus, alpha is a probability about the test’s decision given that the null model is true. It is not a probability about whether the null hypothesis is true. Likewise, power is a probability about the test’s decision given a particular alternative value, not a general probability that the research claim is correct.

For a particular test result, the population proportion is unknown. We therefore do not use alpha to say how likely it is that \(H_0\) is true after seeing the data. A significance test does not calculate that probability. Its p-value answers a different question: assuming \(H_0\) is true, how likely is a result at least as extreme as the one observed? As explained in “Significance Level as the Probability of a Type I Error,” alpha is chosen as a standard for the testing procedure; the p-value is calculated from the sample.

Alpha Is Not the Probability That the Null Is True

A common misconception is, “If \(\alpha=0.05\), there is a 5% chance that the null hypothesis is true.” That reverses the condition. Alpha does not describe the probability of a hypothesis being true. It describes the probability that the test rejects when the null hypothesis is true.

Think about repeating the same testing procedure many times under the same conditions, with the null hypothesis actually true each time. If the procedure has significance level \(\alpha=0.05\), its long-run probability of rejection is 0.05. In many repetitions, about 5% of the tests would reject a true null, although the exact percentage in a finite set of repetitions need not be exactly 5%.

That long-run statement does not tell us whether the null hypothesis is true in any one investigation. Nor does it mean that every rejection has a 5% probability of being wrong. Whether a particular rejection is a Type I error depends on the actual population truth, which the test generally does not reveal.

Worked Example: Interpreting Alpha for a Recycling Survey

A city uses a significance test to assess whether more than 40% of residents regularly recycle. The procedure uses \(\alpha=0.05\). A student says, “There is a 5% probability that the true proportion of residents who recycle is 40%.” Is that a correct interpretation?

Solution: No. The statement treats alpha as the probability that the null hypothesis is true. The null hypothesis is \(H_0:p=0.40\), where \(p\) is the proportion of city residents who regularly recycle. Alpha instead describes the test’s behavior under the condition that this null value is true:

$$ \alpha=P(\text{reject }H_0\mid p=0.40)=0.05 $$

In repeated use of this procedure when the true proportion is 0.40, the probability of rejecting \(H_0\) is 0.05. If the procedure were applied 1,000 times under that condition, the expected number of rejections would be \(1{,}000(0.05)=50\). This is a long-run expectation, not a promise that exactly 50 of those 1,000 tests would reject.

Conclusion: Alpha is the probability of a Type I error for the testing procedure, not the probability that the null hypothesis is true. For a single survey, alpha alone does not tell us the probability that the city’s true recycling proportion equals 0.40.

Power Is Conditional on a Specified Alternative

A second common misconception is, “Power is the probability of rejecting.” That wording is incomplete. A test can have different probabilities of rejection at different true population proportions. To state power, specify the alternative value being treated as true.

Suppose a test assesses \(H_0:p=p_0\), and we want to describe its ability to detect a particular change. We select a value \(p_1\) that makes the null hypothesis false. The probability of failing to reject \(H_0\) when \(p=p_1\) is \(\beta(p_1)\). The probability of rejecting under that same condition is the power at \(p_1\):

$$ \text{Power at }p_1 =P(\text{reject }H_0\mid p=p_1) =1-\beta(p_1) $$

The phrase “at \(p_1\)” matters. A test designed to detect a large difference from \(p_0\) may have high power for that value but lower power for a value closer to \(p_0\). As discussed in “How Effect Size Affects Power,” power generally increases when the true alternative is farther from the null in the direction of the alternative, with the sample size and significance level held fixed.

Power is not the probability that the alternative hypothesis is true. It is also not a guarantee that the test will reject in one particular sample. It describes a long-run probability under the stated alternative condition.

Worked Example: Interpreting Power for a Reminder Program

A school tests whether a text-message reminder increases the proportion of students who submit a form on time, compared with a benchmark of 0.60. For a specified true proportion of \(p_1=0.70\), the test has power 0.80. Find beta at \(p_1=0.70\) and explain both values in context.

Solution: The power is the probability of rejecting \(H_0:p=0.60\) when the true proportion submitting on time is 0.70. The probability of failing to reject under that same true value is beta. Since the two possible test decisions—reject and fail to reject—have probabilities that add to 1 under a fixed true value:

$$ \beta(0.70)=1-\text{power at }0.70=1-0.80=0.20 $$

Interpretation: If the true proportion of students who submit on time is 0.70, this procedure has an 0.80 probability of rejecting \(H_0\) and a 0.20 probability of failing to reject \(H_0\). The 0.20 is the probability of a Type II error at \(p=0.70\). These values do not say there is an 80% chance that the true proportion is 0.70.

Conclusion: State the assumed true value whenever reporting power or beta. “The power is 0.80” is less informative than “The power is 0.80 when the true on-time submission proportion is 0.70.”

“Probability of Rejecting” Needs a Truth Condition

The test procedure has a rejection rule, but the probability that a sample lands in the rejection region changes with the true population value. Under the null value, that probability is alpha. Under a specified false-null value, it is power at that value. If the true value is not specified, “the probability of rejecting” does not identify one particular probability.

The following hypothetical operating characteristics make the distinction visible. Suppose a test of \(H_0:p=0.40\) against \(H_a:p>0.40\) has significance level 0.05, and its calculated power is 0.30 when \(p=0.45\) and 0.80 when \(p=0.55\). These values describe one illustrative procedure; they are not results from an actual study.

Assumed true proportionProbability of rejecting \(H_0\)Meaning
\(p=0.40\)0.05Alpha: probability of rejecting a true null
\(p=0.45\)0.30Power at \(p=0.45\)
\(p=0.55\)0.80Power at \(p=0.55\)

All three entries concern the chance of rejection, but they condition on different true proportions. The first is alpha; the other two are powers. The larger power at \(p=0.55\) reflects that this value is farther above the null value than \(p=0.45\), making it easier for the test to detect the increase.

Worked Example: Comparing Power at Two True Values

Use the hypothetical test summarized above. Explain why it is inaccurate to say “the probability of rejecting \(H_0\) is 0.80” without further qualification. Also find beta at \(p=0.45\) and at \(p=0.55\).

Solution: The rejection probability is not a single number for all possible true proportions. Under \(p=0.40\), the rejection probability is 0.05. Under \(p=0.45\), it is 0.30. Under \(p=0.55\), it is 0.80. Thus, 0.80 is the power only when the true proportion is 0.55.

For each specified value, beta is the probability of failing to reject, so subtract the corresponding power from 1:

$$ \beta(0.45)=1-0.30=0.70 $$
$$ \beta(0.55)=1-0.80=0.20 $$

Interpretation: If the true proportion is 0.45, the test has a 0.70 probability of failing to reject the null. If the true proportion is 0.55, it has a 0.20 probability of failing to reject. The test is more likely to detect the larger increase, but neither power value says what the true proportion actually is.

Do Not Confuse Alpha, Power, and the P-Value

These three quantities all involve probabilities, but they answer different questions. Keeping their conditions and roles separate prevents many interpretation errors.

  • Alpha: Before collecting data, the investigator chooses a significance level. It is the probability that the testing procedure rejects \(H_0\) when \(H_0\) is true.
  • Power: For a specified true value that makes \(H_0\) false, power is the probability that the procedure rejects \(H_0\).
  • P-value: After observing sample data, the p-value is calculated under the assumption that \(H_0\) is true. It measures how unusual the observed result, or a more extreme result, would be under that null model.

Alpha is a chosen decision standard; the p-value is evidence from the observed sample compared with that standard. Power describes how often the procedure would reject under a specified alternative. A p-value is not the probability that \(H_0\) is true, and alpha is not the p-value for every sample. Power is not the probability that a rejection is correct.

The distinction also clarifies why “fail to reject” is not proof of the null. A test can fail to reject when the null is true, which is a correct decision, or when a false null is true, which is a Type II error. How likely the latter is depends on beta at the actual alternative value.

Common Mistakes and AP Exam Tip

  • Reversing the condition in alpha: “Alpha is the probability that \(H_0\) is true” is incorrect. Full-credit wording says alpha is the probability of rejecting a true \(H_0\), or \(P(\text{reject }H_0\mid H_0\text{ is true})\).
  • Leaving the alternative value out of a power interpretation: “The test has 80% power” should be tied to a specified true value, such as “when the population proportion is 0.70.” Power can differ at other false-null values.
  • Calling power the probability of any rejection: State what is assumed true. Under the null, the rejection probability is alpha; under a particular alternative, it is power at that alternative.
  • Calling power the probability that the alternative is true: Power conditions on the alternative being true. It does not estimate the chance that this condition holds in the population.
  • Mixing up power and beta: At the same specified alternative, \(\beta\) is the probability of failing to reject, while power is the probability of rejecting. They add to 1.
  • Confusing alpha with a p-value: Alpha is set as a threshold before the sample result is evaluated. The p-value comes from the observed data under the null model. Compare them to make a test decision; do not describe them as the same quantity.
  • Using “the chance this decision is wrong”: Neither alpha nor power directly gives that probability for an individual test result. They describe conditional long-run behavior under specified truths.

On an AP response, include the condition in your sentence: “If the true proportion is 0.70, the test has a 0.80 probability of rejecting \(H_0\).” For alpha, name the null truth: “If \(H_0\) is true, the test has a 0.05 probability of rejecting it.” Conditional wording makes clear that these probabilities describe the procedure, not certainty about the unknown population.

Key takeaway: Alpha is the probability of rejecting a true null. Power is the probability of rejecting when a specified alternative value is true, and beta is the probability of failing to reject at that same value. None of these is the probability that a hypothesis is true.

Check Your Understanding

For each item, identify the condition attached to the probability and interpret it carefully.

  1. A test uses \(\alpha=0.01\). Write the correct conditional interpretation of this significance level.
  2. A test has power 0.75 when the true population proportion is 0.62. What is beta at \(p=0.62\), and what does it mean in context?
  3. Why is “the test has a 0.75 probability of rejecting” incomplete if the specified true proportion is omitted?
  4. A student says, “The p-value is 0.04, so there is a 4% chance that \(H_0\) is true.” Explain the error.
  5. For the same test, suppose power is 0.40 at one alternative value and 0.85 at a farther value in the direction of \(H_a\). What does this tell you about the test’s rejection probabilities at those two values?