Beta Describes a Missed Effect at a Specified Truth
In “Choosing Alpha Based on Error Consequences,” you saw that a test’s significance level is selected before examining the data. The next question is how likely the test is to miss a real effect. That probability is called beta, written \(\beta\). As in “Type II Error Defined in Context,” a Type II error occurs when a test fails to reject a false null hypothesis.
There is an important qualification: beta is calculated for a particular value of the population parameter that makes the null hypothesis false. The probability of failing to reject \(H_0\) can be different for different true values. So, for a test of a population proportion \(p\), it is useful to write \(\beta(p_1)\), where \(p_1\) is a specified true proportion in the alternative hypothesis.
Alpha and beta refer to different conditions. Alpha is the probability of rejecting \(H_0\) when \(H_0\) is true. Beta is the probability of failing to reject \(H_0\) when a specified alternative is true. Beta is not the probability that \(H_0\) is true, and it is not a single fixed feature of a test unless the alternative value has also been specified.
Use the Rejection Cutoff to Find Beta
Consider a one-proportion z test with \(H_0:p=p_0\) and \(H_a:p>p_0\). For a selected significance level \(\alpha\), the test rejects \(H_0\) when the sample proportion \(\hat{p}\) is sufficiently large. The cutoff comes from the null model: it is the value of \(\hat{p}\) that corresponds to the upper-tail critical z-value.
Here, \(z^*\) is the upper-tail critical value for the chosen \(\alpha\), and \(n\) is the sample size. Once the cutoff is set, imagine repeatedly taking samples when the true proportion is \(p_1\). Under that truth, the sampling distribution of \(\hat{p}\) is centered at \(p_1\), with standard deviation \(\sqrt{p_1(1-p_1)/n}\). Beta is the probability that \(\hat{p}\) falls on the non-rejection side of the cutoff.
The symbol \(\approx\) matters: the calculations below use a normal approximation to the sampling distribution of \(\hat{p}\). For a discrete binomial count, an exact calculation would use the binomial probabilities and the test’s integer rejection rule. The normal approximation makes the relationship between the cutoff, the true proportion, and beta visible.
- The data should come from an appropriate random sample or randomized experiment, and observations should be independent.
- For sampling without replacement, check the 10% condition: the sample size is no more than 10% of the population size.
- Check the Large Counts condition under the null model and under the specified alternative: \(np \ge 10\) and \(n(1-p) \ge 10\). Under \(H_0\), use \(p_0\); to approximate beta at \(p_1\), also check using \(p_1\).
Worked Examples: Beta Depends on the Alternative
Worked Example: A Tutoring Reminder Test
A school plans a one-sided test of whether the attendance proportion among students receiving text reminders exceeds the 0.30 benchmark. The hypotheses are \(H_0:p=0.30\) and \(H_a:p>0.30\). The school will take a random sample of \(n=100\) students who received a text reminder and use \(\alpha=0.05\). Estimate beta if the true proportion is \(p_1=0.35\).
State: We want the probability of failing to reject \(H_0\) when the actual attendance proportion is \(0.35\). This is \(\beta(0.35)\).
Plan and conditions: Use the upper-tail one-proportion z-test cutoff, then find the probability that \(\hat{p}\) is below it under the sampling distribution when \(p=0.35\). Assume the sample is random and the population contains at least 1,000 students, so \(100\le0.10(1000)\). Under the null, \(np_0=100(0.30)=30\) and \(n(1-p_0)=70\), both at least 10. Under the specified alternative, \(np_1=100(0.35)=35\) and \(n(1-p_1)=65\), also both at least 10. The Large Counts condition is satisfied for both calculations.
Do: For \(\alpha=0.05\), use \(z^*=1.645\). The null-model standard deviation is \(\sqrt{0.30(0.70)/100}=\sqrt{0.0021}\approx0.04583\). Thus the rejection cutoff is \(0.30+1.645(0.04583)\approx0.37538\). The test rejects when \(\hat{p}\) is about \(0.37538\) or larger.
When the true proportion is \(0.35\), the sampling distribution of \(\hat{p}\) has standard deviation \(\sqrt{0.35(0.65)/100}=\sqrt{0.002275}\approx0.04770\). Standardizing the cutoff using this alternative distribution gives \(z=(0.37538-0.35)/0.04770\approx0.5322\). Therefore, \(\beta(0.35)\approx P(Z<0.5322)\approx0.7027\), rounded to four decimal places.
Conclude: If the true attendance proportion is \(0.35\), this test has an estimated \(0.7027\), or about 70.27%, chance of failing to reject \(H_0\). That is beta for this specified alternative, not for every possible increase above \(0.30\).
Worked Example: A Larger Increase Is Easier to Detect
Keep the same test plan: \(H_0:p=0.30\), \(H_a:p>0.30\), \(n=100\), and \(\alpha=0.05\). The cutoff remains \(0.37538\). Now estimate beta if the true proportion is \(p_1=0.40\), and compare it with the result when \(p_1=0.35\).
The random-sample and 10% conditions are as before. Under \(p_1=0.40\), the Large Counts checks are \(np_1=100(0.40)=40\) and \(n(1-p_1)=100(0.60)=60\), both at least 10. The normal approximation is reasonable.
At \(p_1=0.40\), the standard deviation of \(\hat{p}\) is \(\sqrt{0.40(0.60)/100}=\sqrt{0.0024}\approx0.04899\). The standardized cutoff is \(z=(0.37538-0.40)/0.04899\approx-0.5025\). So \(\beta(0.40)\approx P(Z<-0.5025)\approx0.3077\), rounded to four decimal places.
With the same test, beta is about \(0.7027\) when the true proportion is \(0.35\), but about \(0.3077\) when it is \(0.40\). The rejection cutoff has not changed. When \(p=0.40\), the sampling distribution is centered farther above that cutoff, so more sample proportions fall in the rejection region. This is why beta depends on the true alternative value.
The result does not mean that every increase above \(0.30\) has beta \(0.3077\). It estimates the chance of a Type II error specifically when the true proportion is \(0.40\), using this sample size, significance level, and test.
Worked Example: A Stricter Alpha Changes Beta
Suppose the school keeps \(H_0:p=0.30\), \(H_a:p>0.30\), and \(n=100\), but chooses \(\alpha=0.01\) instead of \(0.05\). Estimate beta when the true proportion is \(p_1=0.40\). Assume the same random-sampling design and population size of at least 1,000 students. The null and alternative Large Counts checks are \(30,70\) and \(40,60\), respectively, so all are at least 10; the 10% condition is also met.
For an upper-tail test with \(\alpha=0.01\), use \(z^*=2.326\). The null standard deviation is still \(\sqrt{0.30(0.70)/100}\approx0.04583\), so the new cutoff is \(0.30+2.326(0.04583)\approx0.40659\). This cutoff is higher than the \(0.37538\) cutoff for \(\alpha=0.05\), so the test requires a more extreme sample proportion to reject \(H_0\).
Under the specified truth \(p_1=0.40\), the standard deviation remains \(\sqrt{0.40(0.60)/100}\approx0.04899\). Thus \(z=(0.40659-0.40)/0.04899\approx0.1345\), and \(\beta(0.40)\approx P(Z<0.1345)\approx0.5535\), rounded to four decimal places.
For the same true proportion and sample size, the estimated beta was about \(0.3077\) at \(\alpha=0.05\) and is about \(0.5535\) at \(\alpha=0.01\). A smaller alpha makes rejection harder, so failing to reject is more likely when \(p=0.40\). This illustrates the tradeoff discussed in “Choosing Alpha Based on Error Consequences”: limiting Type I error probability can increase the chance of a Type II error for a fixed sample size and specified alternative.
What Beta Does—and Does Not—Tell You
Beta answers a conditional question: if the true proportion were \(p_1\), how often would this planned test fail to reject \(H_0\)? It does not tell you the chance that the observed test result is wrong. It also does not tell you whether the true proportion is \(p_1\). The value \(p_1\) is supplied as a possibility to evaluate, not estimated by beta.
A test with a composite alternative, such as \(H_a:p>0.30\), has many possible false-null values: \(0.31\), \(0.35\), \(0.40\), and others. Each can have a different beta. A value close to \(0.30\) is generally harder to distinguish from the null than a value farther above it. Therefore, stating “the beta is 0.31” is incomplete unless the true alternative value and test setup are clear.
The sample size and alpha matter too. Changing either can move the rejection cutoff or change the spread of the sampling distribution. When comparing beta values, keep track of what is held fixed. For example, the comparison in the second example changed the specified true proportion but kept the test plan the same; the third changed alpha while holding the true proportion and sample size fixed.
Common Mistakes and AP Exam Tip
- Reporting beta without specifying the alternative value: A test of \(p>p_0\) has many possible true values under the alternative. Say “when the true proportion is \(p_1=\dots\)” before giving beta.
- Using the null standard deviation for beta: The null standard deviation is used to find the test’s rejection cutoff. To estimate beta at \(p_1\), standardize that cutoff using \(\sqrt{p_1(1-p_1)/n}\), the spread under the specified alternative.
- Finding the rejection probability instead of the error probability: Beta is the area in the non-rejection region when the null is false. For an upper-tail test, that is the area below the cutoff under the alternative distribution.
- Claiming alpha and beta are interchangeable: Alpha is the Type I error probability under a true null. Beta is a Type II error probability under a specified false null. State the truth condition for each.
- Forgetting to check approximation conditions under the alternative: When using a normal approximation to estimate beta, check the Large Counts condition using \(p_1\) as well as \(p_0\). Also address randomization or random sampling and the 10% condition when relevant.
- Overstating an approximate result: If beta is estimated with a normal approximation, identify it as approximate and give a sensible rounding. A full-credit response connects the probability to failing to reject \(H_0\) in the stated context.
Check Your Understanding
Use the ideas in this tutorial to answer each question. Unless a question says otherwise, think about an upper-tail test for a population proportion.
- In your own words, what does \(\beta(0.38)\) mean if the null hypothesis is \(H_0:p=0.30\)?
- Why can the same test have one beta value when \(p=0.34\) and a different beta value when \(p=0.45\)?
- For an upper-tail test, is beta the probability of landing above or below the rejection cutoff when the specified alternative is true? Explain.
- When estimating beta with a normal approximation at \(p_1\), which proportion is used in the standard deviation of \(\hat{p}\): \(p_0\) or \(p_1\)? Why?
- A test changes from \(\alpha=0.05\) to \(\alpha=0.01\), with the sample size and specified true proportion unchanged. What generally happens to the rejection cutoff and beta? Explain the direction of each change.