Why Distance from the Null Matters
In “How Sample Size Affects Power” and “How Significance Level Affects Power and Type II Error,” we changed the sample size or \(\alpha\) while holding other features of a test fixed. Now we hold those features fixed and compare different possible true proportions. The key question is whether a true proportion is close to or far from the null value in the direction of the alternative.
For a one-proportion test, the effect size for a specified alternative is the distance between the true proportion \(p_1\) and the null value \(p_0\). In this tutorial, we focus on alternatives in the direction stated by \(H_a\). When \(H_a:p>p_0\), for example, a true proportion just above \(p_0\) is a smaller effect than a true proportion farther above \(p_0\).
Power is the probability of rejecting \(H_0\) when a specified alternative value is true. As in the earlier tutorials on beta and power, this is a conditional probability: it describes the test’s long-run behavior assuming a particular true proportion. It is not the probability that \(H_a\) is true.
For an upper-tail test, the null model sets the approximate cutoff for \(\hat{p}\). To estimate power at a particular \(p_1\), use the sampling distribution of \(\hat{p}\) centered at that \(p_1\) to find the probability of reaching or exceeding the cutoff. Thus, the null model determines the rejection rule, while the alternative model determines the probability of detecting the effect.
For a lower-tail test, the cutoff is below \(p_0\), and estimated power is the probability of falling at or below that cutoff under the specified alternative. The same principle applies: use the null model to locate the rejection boundary, then use the specified true proportion to calculate the probability of crossing it.
Comparing a Close and a Far Alternative
Worked Example: Two Possible Increases in Tool Use
A community garden is evaluating a digital tool for scheduling shared equipment. The current proportion of members who use the tool is represented by \(p_0=0.40\). A follow-up test asks whether the proportion has increased, using \(H_0:p=0.40\), \(H_a:p>0.40\), \(n=160\), and \(\alpha=0.05\). Compare estimated power if the true proportion is \(p_1=0.44\) with estimated power if it is \(p_1=0.52\).
State: The effect sizes are \(0.44-0.40=0.04\) and \(0.52-0.40=0.12\). We want to compare the probability of rejecting \(H_0\) under each specified truth, with the hypotheses, sample size, alpha, and test direction held fixed.
Plan and conditions: Use an upper-tail normal approximation. Assume the 160 members are selected using an appropriate random process. If the sample is drawn without replacement, suppose the garden has at least 1,600 members, so \(160\leq0.10(1600)\); the 10% condition is met. Under the null, the Large Counts checks are \(np_0=160(0.40)=64\) and \(n(1-p_0)=160(0.60)=96\), both at least 10. For \(p_1=0.44\), the counts are \(160(0.44)=70.4\) and \(160(0.56)=89.6\), both at least 10. For \(p_1=0.52\), they are \(83.2\) and \(76.8\), also both at least 10. These conditions support the normal approximations for both power estimates.
Do: The null standard deviation of \(\hat{p}\) is \(\sqrt{0.40(0.60)/160}=\sqrt{0.0015}\approx0.0387298\). For an upper-tail test with \(\alpha=0.05\), \(z^*\approx1.6449\). The approximate rejection cutoff is \(0.40+1.6449(0.0387298)\approx0.4637049\).
If \(p_1=0.44\), the alternative standard deviation is \(\sqrt{0.44(0.56)/160}=\sqrt{0.00154}\approx0.0392428\). Standardizing the cutoff under this alternative gives \((0.4637049-0.44)/0.0392428\approx0.6041\). Therefore, estimated power is \(P(Z\geq0.6041)\approx0.2729\), rounded to four decimal places.
If \(p_1=0.52\), the alternative standard deviation is \(\sqrt{0.52(0.48)/160}=\sqrt{0.00156}\approx0.0394968\). The standardized cutoff is \((0.4637049-0.52)/0.0394968\approx-1.4253\). Thus estimated power is \(P(Z\geq-1.4253)\approx0.9230\), rounded to four decimal places.
Conclude: If the true proportion of garden members using the tool is \(0.44\), the test has estimated power of about \(0.2729\). If the true proportion is \(0.52\), its estimated power is about \(0.9230\). The farther alternative is much more likely to produce a sample proportion in the rejection region, so this test is more likely to detect that larger increase.
What the Sampling Distributions Show
The cutoff in the garden example is based on the null value \(p_0=0.40\), and it stays the same for both power calculations. Under \(p_1=0.44\), the sampling distribution of \(\hat{p}\) is centered at \(0.44\), fairly close to the cutoff of about \(0.4637\). Many sample proportions under that alternative fall below the cutoff, so power is modest.
Under \(p_1=0.52\), the sampling distribution is centered farther above the same cutoff. A much larger fraction of sample proportions exceed it, which explains the higher estimated power. In each case, the calculation measures an area under the alternative sampling distribution—not the null distribution.
For an upper-tail test, a true proportion closer to \(p_0\) in the direction of \(H_a\) generally means a greater chance of failing to reject \(H_0\), so \(\beta(p_1)\) is larger and power \(1-\beta(p_1)\) is smaller. A true proportion farther in that direction generally has smaller \(\beta(p_1)\) and greater power. These comparisons refer to specified alternatives; beta and power are not single values that apply to every possible false null value.
The effect-size comparison is meaningful only when the test features are held fixed. If sample size, alpha, or test direction also changes, a change in power cannot be attributed only to the distance between \(p_1\) and \(p_0\). Moreover, a proportion far from \(p_0\) in the direction opposite to \(H_a\) does not imply greater power for rejecting \(H_0\) in favor of that stated alternative.
Comparing Two More Increases
Worked Example: A Larger Effect in a Scheduling Survey
A city is testing whether a new online system increases the proportion of residents who book a bulky-item pickup online. The null value is \(p_0=0.30\), and the test is \(H_0:p=0.30\) versus \(H_a:p>0.30\), with \(n=100\) and \(\alpha=0.05\). Compare estimated power at \(p_1=0.40\) and \(p_1=0.50\).
Plan and conditions: Assume an appropriate random sample of residents. If sampling without replacement, suppose there are at least 1,000 residents in the population, so \(100\leq0.10(1000)\). Under the null, \(np_0=30\) and \(n(1-p_0)=70\), both at least 10. Under \(p_1=0.40\), the counts are 40 and 60; under \(p_1=0.50\), they are 50 and 50. All satisfy the Large Counts condition.
Do: The null standard deviation is \(\sqrt{0.30(0.70)/100}=\sqrt{0.0021}\approx0.0458258\). The upper-tail critical value is \(z^*\approx1.6449\), so the common cutoff is \(0.30+1.6449(0.0458258)\approx0.375376\).
At \(p_1=0.40\), the alternative standard deviation is \(\sqrt{0.40(0.60)/100}=\sqrt{0.0024}\approx0.0489898\). The standardized cutoff is \((0.375376-0.40)/0.0489898\approx-0.5027\). The estimated power is \(P(Z\geq-0.5027)\approx0.6924\).
At \(p_1=0.50\), the alternative standard deviation is \(\sqrt{0.50(0.50)/100}=\sqrt{0.0025}=0.0500000\). The standardized cutoff is \((0.375376-0.50)/0.0500000\approx-2.4925\). Estimated power is \(P(Z\geq-2.4925)\approx0.9937\). These probabilities are rounded to four decimal places.
Conclude: With the same test and sample size, estimated power is about \(0.6924\) for a true booking proportion of \(0.40\), but about \(0.9937\) if the true proportion is \(0.50\). The larger increase is farther from \(p_0=0.30\), so the sample proportion is much more likely to cross the same rejection cutoff.
The Direction of the Alternative Still Matters
“Farther from the null” must be understood relative to the alternative hypothesis. For a lower-tail test, the relevant alternatives are below \(p_0\), and the rejection region is below a null-based cutoff. The following example applies the same reasoning in the opposite direction.
Worked Example: Two Possible Decreases in a Recycling Rate
A school is checking whether a new bin arrangement reduced the proportion of students who sort recyclable bottles correctly. The null value is \(p_0=0.60\), with \(H_0:p=0.60\), \(H_a:p<0.60\), \(n=120\), and \(\alpha=0.05\). Compare estimated power at \(p_1=0.55\) and \(p_1=0.40\).
Plan and conditions: Assume the students are chosen through an appropriate random process, and if the sample is taken without replacement, assume the relevant population is at least 1,200 students; then \(120\leq0.10(1200)\). Under \(p_0=0.60\), \(np_0=72\) and \(n(1-p_0)=48\). At \(p_1=0.55\), the corresponding values are 66 and 54. At \(p_1=0.40\), they are 48 and 72. All are at least 10, satisfying the Large Counts condition.
Do: The null standard deviation is \(\sqrt{0.60(0.40)/120}=\sqrt{0.002}\approx0.0447214\). For a lower-tail test at \(\alpha=0.05\), the critical z-value is approximately \(-1.6449\). Therefore, the cutoff is \(0.60-1.6449(0.0447214)\approx0.526440\).
At \(p_1=0.55\), the alternative standard deviation is \(\sqrt{0.55(0.45)/120}=\sqrt{0.0020625}\approx0.0454142\). The standardized cutoff is \((0.526440-0.55)/0.0454142\approx-0.5188\). The estimated power is the lower-tail area \(P(Z\leq-0.5188)\approx0.3019\).
At \(p_1=0.40\), the alternative standard deviation is \(\sqrt{0.40(0.60)/120}=\sqrt{0.002}\approx0.0447214\). The standardized cutoff is \((0.526440-0.40)/0.0447214\approx2.8273\). Estimated power is \(P(Z\leq2.828)\approx0.9977\), rounded to four decimal places.
Conclude: The test has estimated power of about \(0.3019\) if the true sorting proportion is \(0.55\), and about \(0.9977\) if it is \(0.40\). The second alternative is farther below the null value, in the direction of the lower-tail test, and is therefore much easier for this test to detect.
Common Mistakes and AP Exam Tip
- Using the null distribution for the power area: Use \(p_0\) to establish the rejection cutoff, but use the specified \(p_1\) to calculate the probability of crossing it.
- Reporting power without naming the alternative: Power depends on the true value assumed. State the value of \(p_1\) and interpret the probability in context.
- Calling every distant alternative easier to detect: The distance must be in the direction of the stated alternative. A lower-tail test is designed to detect values below \(p_0\), not increases above it.
- Changing several features in a comparison: To isolate effect size, hold the sample size, alpha, null value, and direction of the test fixed.
- Describing an approximation as exact: The sample proportion is discrete, while the normal model is continuous. Identify normal-approximation power values as estimates.
- Confusing power with evidence from one sample: Power describes the chance of rejection over repeated samples when a specified true proportion holds; it is not a conclusion about the truth based on one observed sample.
For a full-credit comparison, identify what stays fixed, give both specified true proportions, and compare the resulting estimated powers in context. Explain that the more distant alternative in the direction of \(H_a\) is more likely to produce a sample proportion in the rejection region.
Check Your Understanding
Use the relationship between effect size and power to answer each question.
- For an upper-tail test of \(H_0:p=0.25\) versus \(H_a:p>0.25\), which has the larger effect size: \(p_1=0.28\) or \(p_1=0.40\)?
- When estimating power for a specified \(p_1\), which proportion determines the rejection cutoff, and which determines the sampling distribution used for the power probability?
- Why should sample size, alpha, and test direction stay fixed when comparing power for two effect sizes?
- For a lower-tail test, explain why a true proportion farther below \(p_0\) generally has greater power than one just below \(p_0\).
- In context, what does estimated power of \(0.80\) at a specified true proportion mean?