From a Rejection Rule to Power
In “Power of a Test Defined” and “How Effect Size Affects Power,” we described power as the probability of rejecting \(H_0\) when a specified alternative value is true. Now we will calculate that probability for a one-proportion z test. The key is to find the test’s rejection region first, then determine how often sample proportions from the specified alternative would fall in that region.
For an upper-tail test, the rejection region consists of sufficiently large values of \(\hat{p}\). For a lower-tail test, it consists of sufficiently small values. For a two-sided test, it includes values in either tail. The rejection boundary is set using the null proportion \(p_0\) and the significance level \(\alpha\). To calculate power, however, use the sampling distribution centered at the specified true proportion \(p_1\).
For a one-proportion test, the approximate standard deviation of \(\hat{p}\) when the true proportion is \(p\) is \(\sqrt{p(1-p)/n}\). The cutoff uses \(p_0\) in this formula; the power calculation uses \(p_1\). In the examples, we use normalcdf to calculate the area under the alternative sampling distribution. Because \(\hat{p}\) is discrete but the normal model is continuous, this gives an approximate power.
For a lower-tail test, use the lower-tail critical value in the cutoff, or subtract the corresponding positive critical value. For a two-sided test, find both cutoffs and add the two tail probabilities under \(p_1\). The test’s direction determines which areas count as rejection.
Worked Example: Power for an Increase in Helmet Use
Worked Example: Power for an Increase in Helmet Use
A community program aims to increase the proportion of bicycle commuters who wear a helmet. A test uses \(H_0:p=0.30\) versus \(H_a:p>0.30\), with \(n=200\) commuters and \(\alpha=0.05\). Estimate power if the true proportion wearing a helmet is \(p_1=0.40\).
State: We want the probability that this test rejects \(H_0\), assuming the true helmet-use proportion is \(0.40\). This is the test’s estimated power at \(p_1=0.40\).
Plan and conditions: Use an upper-tail normal approximation to the sampling distribution of \(\hat{p}\). Assume the commuters are selected through an appropriate random process. If they are sampled without replacement, suppose the population contains at least 2,000 bicycle commuters; then \(200\leq0.10(2000)\), so the 10% condition is met. Under the null, the Large Counts condition is met because \(np_0=200(0.30)=60\) and \(n(1-p_0)=200(0.70)=140\), both at least 10. Under the specified alternative, \(np_1=200(0.40)=80\) and \(n(1-p_1)=200(0.60)=120\), also both at least 10. These checks support the normal approximations used for the cutoff and power.
Do: The upper-tail critical value for \(\alpha=0.05\) is \(z^*\approx1.6449\). Under the null, the standard deviation of \(\hat{p}\) is \(\sqrt{0.30(0.70)/200}=\sqrt{0.00105}\approx0.0324037\). Therefore, the approximate rejection cutoff is \(0.30+1.6449(0.0324037)\approx0.353299\). The test rejects for sample proportions at or above about \(0.3533\).
Under \(p_1=0.40\), the standard deviation of \(\hat{p}\) is \(\sqrt{0.40(0.60)/200}=\sqrt{0.0012}\approx0.0346410\). Use this alternative distribution—not the null distribution—to find the probability of reaching the rejection region:
As a check, standardizing the cutoff under \(p_1=0.40\) gives \((0.353299-0.40)/0.0346410\approx-1.3481\). The area above a z-score of \(-1.3481\) is about \(0.9112\), agreeing with the calculator result.
Conclude: If the true proportion of bicycle commuters who wear a helmet is \(0.40\), this test has estimated power of about \(0.9112\). In repeated samples of 200 commuters under that specified truth, the test would reject \(H_0\) about 91.12% of the time. This is a conditional probability, not the probability that the true proportion is \(0.40\).
Lower-Tail Tests Use the Lower Rejection Region
The method is the same for a lower-tail test, but the rejection region is below the null-based cutoff. The following example includes the entire lower-tail area under the alternative distribution.
Worked Example: Power for a Decrease in Water-Filter Defects
A manufacturer tests whether a revised process lowers the proportion of water-filter cartridges that fail a quality check. The hypotheses are \(H_0:p=0.50\) and \(H_a:p<0.50\), with \(n=150\) cartridges and \(\alpha=0.10\). Estimate power if the true defect proportion is \(p_1=0.40\).
Plan and conditions: Assume the cartridges are selected through an appropriate random process. If they are sampled without replacement, suppose the production population contains at least 1,500 cartridges, so \(150\leq0.10(1500)\); the 10% condition is met. Under the null, \(np_0=75\) and \(n(1-p_0)=75\), both at least 10. Under \(p_1=0.40\), the counts are \(np_1=60\) and \(n(1-p_1)=90\), also both at least 10. Thus the normal approximation is supported under both proportions.
Do: For a lower-tail test with \(\alpha=0.10\), the critical z-value is approximately \(-1.2816\). The null standard deviation is \(\sqrt{0.50(0.50)/150}=\sqrt{0.0016667}\approx0.0408248\). The cutoff is \(0.50-1.2816(0.0408248)\approx0.447681\). The test rejects when \(\hat{p}\) is at or below about \(0.4477\).
At \(p_1=0.40\), the standard deviation is \(\sqrt{0.40(0.60)/150}=\sqrt{0.0016}=0.0400000\). Therefore, \(\operatorname{normalcdf}(-1\mathrm{E}99,0.447681,0.40,0.0400000)\approx0.8834\). Equivalently, the standardized cutoff is \((0.447681-0.40)/0.0400000\approx1.1920\), and the area to its left is about \(0.8834\).
Conclude: If the true cartridge defect proportion is \(0.40\), the test’s estimated power is about \(0.8834\). In repeated samples of 150 cartridges under that alternative, the test would reject the null hypothesis about 88.34% of the time.
Two-Sided Tests Have Two Rejection Tails
For a two-sided alternative, \(H_a:p\ne p_0\), the test rejects for sample proportions that are unusually low or unusually high relative to \(p_0\). Estimate power by calculating the probability of each tail under the specified alternative and adding them. The two probabilities may not be equally large, especially when \(p_1\) is on one side of \(p_0\).
Worked Example: Power for a Change in App Use
A library tests whether the proportion of visitors who use its room-booking app has changed from \(0.50\). The test is \(H_0:p=0.50\) versus \(H_a:p\ne0.50\), with \(n=100\) visitors and \(\alpha=0.05\). Estimate power if the true proportion using the app is \(p_1=0.65\).
Plan and conditions: Assume visitors are selected through an appropriate random process. If sampling without replacement, suppose the relevant population contains at least 1,000 visitors, so \(100\leq0.10(1000)\). Under the null, \(np_0=50\) and \(n(1-p_0)=50\); under \(p_1=0.65\), the counts are 65 and 35. All are at least 10, so the Large Counts condition is met under both values.
Do: For a two-sided test with \(\alpha=0.05\), the critical z-values are approximately \(-1.96\) and \(1.96\). Under the null, the standard deviation is \(\sqrt{0.50(0.50)/100}=0.0500000\). The lower and upper rejection cutoffs are \(0.50-1.96(0.0500000)=0.402000\) and \(0.50+1.96(0.0500000)=0.598000\). Thus the rejection region is approximately \(\hat{p}\leq0.4020\) or \(\hat{p}\geq0.5980\).
Under \(p_1=0.65\), the standard deviation is \(\sqrt{0.65(0.35)/100}=\sqrt{0.002275}\approx0.0476970\). Add the alternative probabilities in both rejection tails:
The lower-tail probability is extremely small: the standardized lower cutoff is \((0.402000-0.65)/0.0476970\approx-5.1995\). The upper-tail standardized cutoff is \((0.598000-0.65)/0.0476970\approx-1.0902\), giving an upper-tail probability of about \(0.8622\). The tiny lower-tail area rounds to \(0.0000\) at four decimal places; it is included in the total before rounding.
Conclude: If the true proportion of library visitors using the app is \(0.65\), the two-sided test has estimated power of about \(0.8622\). Most of that power comes from the upper rejection tail, because this specified alternative is above the null value.
Common Mistakes and AP Exam Tip
- Using \(p_1\) to set the rejection cutoff: The cutoff comes from \(p_0\), the test’s null model. Use \(p_1\) only for the sampling distribution that supplies the power area.
- Using the null standard deviation for the power area: Calculate \(\sqrt{p_1(1-p_1)/n}\) when finding power for a specified alternative. The standard deviation under \(p_0\) is used to locate the cutoff.
- Taking the wrong tail: An upper-tail test uses the area above its cutoff; a lower-tail test uses the area below it. A two-sided test requires both tails.
- Reporting power without specifying the truth assumed: Power is calculated for a particular \(p_1\). Name that value and interpret the result conditionally, in context.
- Calling the normal approximation exact: The normalcdf result is an estimate because the sample proportion is discrete. Keep enough digits in the cutoff and standard deviation when entering calculator values, then round the final probability consistently.
- Skipping conditions: Check randomization or random sampling, the 10% condition when sampling without replacement, and the Large Counts condition under both \(p_0\) and the specified \(p_1\).
For a full-credit response, identify the hypotheses and specified alternative, show how the rejection cutoff comes from the null model, check the conditions, and use the correct alternative distribution and tail area. Finish with a conditional interpretation in context.
Check Your Understanding
Use the rejection-region method to answer each question.
- For an upper-tail test, which proportion determines the rejection cutoff: \(p_0\) or the specified \(p_1\)?
- For an upper-tail test with cutoff \(0.32\), \(p_1=0.38\), and alternative standard deviation \(0.04\), which normalcdf area represents estimated power?
- Why should you check the Large Counts condition under the specified alternative as well as under the null?
- For a two-sided test, what two areas must be added to estimate power?
- In context, what does estimated power of \(0.75\) at a specified true proportion mean?