Why Sample Size Matters for Power
In “Power of a Test Defined,” you learned that power is the probability of rejecting \(H_0\) when a specified alternative value is true. Now we will compare that probability for tests that differ only in sample size. The central idea is that a larger sample generally makes the sample proportion vary less from sample to sample, so a real departure from the null is easier to distinguish from ordinary sampling variation.
We will use a one-proportion z test with \(H_0:p=0.30\), \(H_a:p>0.30\), and a specified alternative of \(p_1=0.40\). The significance level will remain \(\alpha=0.05\). The sample size changes, but the hypotheses, direction of the test, alpha and specified truth stay fixed. This makes the power comparison meaningful.
For this upper-tail test, the approximate rejection cutoff for \(\hat{p}\) is \(p_0+z^*\sqrt{p_0(1-p_0)/n}\), where \(p_0=0.30\) and \(z^*=1.645\) for \(\alpha=0.05\). To estimate power at \(p_1=0.40\), find the probability that \(\hat{p}\) is at least that cutoff using the sampling distribution centered at \(p_1\), with standard deviation \(\sqrt{p_1(1-p_1)/n}\). The cutoff is based on the null model; the probability for power is calculated under the specified alternative.
The calculations below use the normal approximation and the conventional critical value \(z^*=1.645\). Because a sample proportion is discrete, this is an estimate of power; an exact binomial calculation can differ slightly. The comparison is still useful for seeing the effect of sample size.
Compare Power While Holding Alpha Fixed
Worked Example: Power With 100 Observations
A community group plans to test whether more than 30% of households use a curbside compost service. It will use \(H_0:p=0.30\), \(H_a:p>0.30\), and \(\alpha=0.05\). Estimate the power when \(n=100\) if the true proportion is \(p_1=0.40\).
State: We want the probability that this test rejects \(H_0:p=0.30\) if the true proportion of households using the service is \(0.40\). This is the approximate power at \(p_1=0.40\).
Plan and conditions: Use the null distribution to set the upper-tail rejection cutoff, then calculate the probability of reaching that cutoff under the sampling distribution at \(p_1=0.40\). Assume the households are selected using an appropriate random sample. If sampling without replacement, suppose there are at least 1,000 households in the population; then \(100\leq0.10(1000)\), so the 10% condition is met. Under \(H_0\), the Large Counts checks are \(np_0=100(0.30)=30\) and \(n(1-p_0)=100(0.70)=70\), both at least 10. Under the specified alternative, they are \(np_1=100(0.40)=40\) and \(n(1-p_1)=100(0.60)=60\), also both at least 10. The normal approximation is reasonable under these assumptions.
Do: Under the null, the standard deviation of \(\hat{p}\) is \(\sqrt{0.30(0.70)/100}=\sqrt{0.0021}\approx0.04583\). The approximate rejection cutoff is \(0.30+1.645(0.04583)\approx0.37538\). Thus the test rejects for sample proportions of about \(0.37538\) or greater.
When \(p=0.40\), the standard deviation of \(\hat{p}\) is \(\sqrt{0.40(0.60)/100}=\sqrt{0.0024}\approx0.04899\). Standardizing the cutoff under this alternative gives \(z=(0.37538-0.40)/0.04899\approx-0.5025\). Therefore, the estimated power is \(P(Z\geq-0.5025)\approx0.6923\), rounded to four decimal places.
Conclude: If 40% of households use the compost service, the planned test with 100 observations has an estimated power of about \(0.6923\), or 69.23%. In repeated samples under that specified truth, it would reject the 30% null about 69.23% of the time.
Worked Example: Increasing the Sample to 400
Keep the compost-service hypotheses, significance level and specified alternative the same, but increase the sample size to \(n=400\). Estimate the power at \(p_1=0.40\), and compare it with the estimate for \(n=100\).
Plan and conditions: As before, assume an appropriate random sample. If the sample is taken without replacement, suppose the population has at least 4,000 households, giving \(400\leq0.10(4000)\) and meeting the 10% condition. Under the null, \(np_0=400(0.30)=120\) and \(n(1-p_0)=400(0.70)=280\), both at least 10. Under the alternative, \(np_1=400(0.40)=160\) and \(n(1-p_1)=400(0.60)=240\), also both at least 10. Thus the normal approximation is reasonable for this comparison as well.
Do: Under the null, the standard deviation is \(\sqrt{0.30(0.70)/400}=\sqrt{0.000525}\approx0.02291\). The rejection cutoff is \(0.30+1.645(0.02291)\approx0.33769\). Notice that this cutoff is closer to \(0.30\) than the cutoff for \(n=100\).
Under \(p=0.40\), the standard deviation is \(\sqrt{0.40(0.60)/400}=\sqrt{0.0006}\approx0.02449\). The standardized cutoff is \(z=(0.33769-0.40)/0.02449\approx-2.5437\). Therefore, \(P(Z\geq-2.5437)\approx0.9945\), rounded to four decimal places.
Conclude: With 400 observations, the test has estimated power of about \(0.9945\), or 99.45%, when the true proportion is \(0.40\). The hypotheses, alpha and specified true proportion did not change; the larger sample increased estimated power from about \(0.6923\) to \(0.9945\).
What Changes—and What Stays Fixed
The null standard error decreases as \(n\) increases, so the null-based cutoff \(p_0+z^*\sqrt{p_0(1-p_0)/n}\) moves toward \(p_0\). Here the cutoff moved from about \(0.37538\) to \(0.33769\). At the same time, the alternative standard error also decreases. The alternative distribution is centered at \(0.40\), so its values cluster more tightly around \(0.40\). Together, these changes make it much more likely that \(\hat{p}\) crosses the rejection cutoff.
Alpha stays fixed in this comparison. It controls the probability of rejecting a true null, so changing \(n\) does not mean we intentionally choose a more permissive significance level. Instead, the cutoff is recalculated for each sample size to keep the null rejection probability near \(0.05\) under the normal approximation. The cutoff shifts because the sampling variability changes.
A useful way to see the increase is to compare the cutoff’s position in standard-deviation units under the alternative. For \(n=100\), the cutoff is about \(0.5025\) standard deviations below \(p_1=0.40\). For \(n=400\), it is about \(2.5437\) standard deviations below \(p_1\). The larger standardized distance leaves less of the alternative distribution below the cutoff, so beta decreases and power, \(1-\beta\), increases.
Worked Example: An Intermediate Sample Size
For the same test, use \(n=225\), with \(H_0:p=0.30\), \(H_a:p>0.30\), \(\alpha=0.05\), and \(p_1=0.40\). Estimate power and place it alongside the estimates for \(n=100\) and \(n=400\).
Plan and conditions: Assume the sample is random and, if it is drawn without replacement, that the population has at least 2,250 households, so \(225\leq0.10(2250)\). Under the null, \(np_0=225(0.30)=67.5\) and \(n(1-p_0)=225(0.70)=157.5\), both at least 10. Under the alternative, \(np_1=225(0.40)=90\) and \(n(1-p_1)=225(0.60)=135\), both at least 10. The Large Counts and sampling conditions support the normal approximation.
Do: The null standard deviation is \(\sqrt{0.30(0.70)/225}=\sqrt{0.0009333}\approx0.03055\). Thus the cutoff is \(0.30+1.645(0.03055)\approx0.35026\). The alternative standard deviation is \(\sqrt{0.40(0.60)/225}=\sqrt{0.0010667}\approx0.03266\). Standardizing the cutoff gives \(z=(0.35026-0.40)/0.03266\approx-1.5231\). So estimated power is \(P(Z\geq-1.5231)\approx0.9361\), rounded to four decimal places.
Conclude: Under the same test plan and specified truth, estimated power rises from about \(0.6923\) at \(n=100\), to \(0.9361\) at \(n=225\), to \(0.9945\) at \(n=400\). The progression illustrates how increasing the sample size can make this specified difference easier to detect.
| Sample size \(n\) | Approximate cutoff | Estimated power at \(p_1=0.40\) |
|---|---|---|
| 100 | 0.37538 | 0.6923 |
| 225 | 0.35026 | 0.9361 |
| 400 | 0.33769 | 0.9945 |
How to Make a Fair Power Comparison
To isolate the effect of sample size, hold the significance level, hypotheses and specified alternative value constant. If you also change \(p_1\), then you are changing how far the truth is from the null, which affects power too. If you change alpha, you are changing the rejection rule as well. A comparison is clearest when it says exactly what stayed fixed.
The trend is not a promise that every larger study will detect an effect. Power is a probability over repeated samples, not a guarantee for one sample. Nor does a high power value mean the alternative is likely to be true. It says that the planned test is likely to reject the null if the specified alternative is in fact the truth.
Power also depends on the size of the departure from the null. A true proportion of \(0.40\) is farther from \(0.30\) than a true proportion of \(0.34\), so the same sample size will generally have more power for \(p_1=0.40\). When interpreting a power calculation, name the alternative value and the population question, as in the earlier tutorial on power.
Common Mistakes and AP Exam Tip
- Claiming that alpha increases with sample size: Alpha is held at the chosen significance level. The cutoff changes with \(n\) so that the test keeps that nominal Type I error rate; it is not correct to say that a larger sample automatically uses a larger alpha.
- Using the null distribution to calculate power: The null distribution determines the rejection cutoff. Calculate the probability of crossing that cutoff using the sampling distribution at the specified alternative \(p_1\).
- Changing several things at once: A comparison that changes \(n\), alpha and \(p_1\) cannot show the effect of sample size alone. State what is held constant.
- Leaving the alternative unspecified: Power is tied to a particular true proportion. A full-credit interpretation says, “If the true proportion is \(0.40\), this test will reject the 30% null in about this fraction of repeated samples.”
- Forgetting the approximation conditions: For a normal-approximation power estimate, identify the random sampling or random assignment basis, check the 10% condition when sampling without replacement, and check Large Counts under both the null and specified alternative.
- Presenting an estimate as an exact probability: The normal model approximates the discrete sampling distribution of \(\hat{p}\). Call the result approximate and keep rounding consistent through the cutoff, standardized value and probability.
Check Your Understanding
Use the idea of power as a conditional probability to answer each question.
- Why must the hypotheses, alpha and specified alternative remain fixed to isolate the effect of sample size?
- For an upper-tail one-proportion test, what distribution determines the rejection cutoff, and what distribution is used to estimate power?
- As \(n\) increases with the other test details fixed, what happens to the null standard error and the approximate rejection cutoff?
- A test has estimated power \(0.90\) at a specified true proportion. Give an appropriate interpretation of that value in context.
- Why should a normal-approximation power calculation be described as approximate rather than exact?