Power Is the Chance of Rejecting a False Null
In “Probability of a Type II Error and Beta,” you learned that \(\beta(p_1)\) is the probability of failing to reject \(H_0\) when the true population proportion is the specified value \(p_1\). Power describes the complementary outcome: the test rejects \(H_0\) when that specified alternative is true.
The phrase “chance to detect a real effect” is a useful shorthand. More precisely, power is the chance that a particular testing procedure will reject its null hypothesis if the population parameter is at the specified alternative value. In context, rejecting \(H_0\) means the test provides evidence against the null and in the direction described by the alternative hypothesis. It does not guarantee that the test will reject \(H_0\) in every sample, even when a real difference exists.
Power is conditional on a specified truth, just as beta is. If the alternative hypothesis includes many possible values, such as \(H_a:p>0.40\), there is not one power value for the entire alternative. The test can have one power when \(p=0.46\), another when \(p=0.52\), and so on. Each value corresponds to the probability of rejecting \(H_0\) at that particular truth.
Power and Beta Are Complementary Outcomes
For a fixed true value \(p_1\), the planned test either rejects \(H_0\) or fails to reject \(H_0\). These are complementary outcomes. If \(p_1\) is true and makes \(H_0\) false, failing to reject is a Type II error; rejecting is the outcome counted by power. Therefore, their probabilities add to 1.
When estimating power with a normal approximation for a one-proportion test, you can find beta as the probability that the sample proportion falls in the non-rejection region under the specified alternative. Subtract that probability from 1 to obtain power. Equivalently, find the probability that the sample proportion falls in the rejection region under the alternative distribution.
Keep the two distributions in their proper roles. The null distribution determines the test’s rejection cutoff. The sampling distribution at \(p_1\) determines how likely the sample proportion is to land on either side of that cutoff when \(p_1\) is true. This is the same distinction used in the previous tutorial when estimating beta.
Worked Examples: Interpreting and Calculating Power
Worked Example: Detecting Bus Use Above a Benchmark
A city analyst wants to test whether more than 40% of residents’ weekend trips involve a bus. The analyst plans a one-proportion z test with \(H_0:p=0.40\), \(H_a:p>0.40\), \(n=120\), and \(\alpha=0.05\). Estimate the test’s power if the true proportion is \(p_1=0.52\).
State: We want the probability that this test rejects \(H_0:p=0.40\) when the true proportion of weekend trips involving a bus is \(0.52\). This is \(\text{Power}(0.52)=1-\beta(0.52)\).
Plan and conditions: Use the upper-tail rejection cutoff for the test, then find the probability of landing in the rejection region under the sampling distribution when \(p=0.52\). Assume the trips are selected through an appropriate random sample and that each sampled trip is recorded independently. If sampling without replacement, suppose the population includes at least 1,200 trips, so \(120\le0.10(1200)\) and the 10% condition is met. Under \(H_0\), the Large Counts checks are \(np_0=120(0.40)=48\) and \(n(1-p_0)=120(0.60)=72\), both at least 10. At the specified alternative, they are \(np_1=120(0.52)=62.4\) and \(n(1-p_1)=120(0.48)=57.6\), also both at least 10. The normal approximation is reasonable under these assumptions.
Do: For an upper-tail test with \(\alpha=0.05\), use \(z^*=1.645\). Under the null, the standard deviation of \(\hat{p}\) is \(\sqrt{0.40(0.60)/120}=\sqrt{0.002}\approx0.04472\). So the rejection cutoff is \(0.40+1.645(0.04472)\approx0.47357\). The test rejects for sample proportions of about \(0.47357\) or greater.
When \(p=0.52\), the standard deviation of \(\hat{p}\) is \(\sqrt{0.52(0.48)/120}=\sqrt{0.00208}\approx0.04561\). Standardizing the cutoff using this alternative distribution gives \(z=(0.47357-0.52)/0.04561\approx-1.0181\). Thus the estimated probability of failing to reject is \(\beta(0.52)\approx P(Z<-1.0181)\approx0.1544\). Therefore, \(\text{Power}(0.52)\approx1-0.1544=0.8456\), rounded to four decimal places.
Conclude: If 52% of weekend trips in the population involve a bus, this test has an estimated 0.8456, or about 84.56%, chance of rejecting the 40% null hypothesis. That is the test’s approximate power at \(p_1=0.52\), not a guarantee about the result from any one sample.
Worked Example: A Smaller Difference Is Harder to Detect
Keep the same bus-use test: \(H_0:p=0.40\), \(H_a:p>0.40\), \(n=120\), and \(\alpha=0.05\). The rejection cutoff remains about \(0.47357\). Estimate power if the true proportion is \(p_1=0.46\).
The random-sampling and 10% conditions remain as described in the first example. Under \(p_1=0.46\), the Large Counts checks are \(np_1=120(0.46)=55.2\) and \(n(1-p_1)=120(0.54)=64.8\), both at least 10. The null checks are still 48 and 72, so the normal approximation conditions are met.
Under \(p=0.46\), the standard deviation of \(\hat{p}\) is \(\sqrt{0.46(0.54)/120}=\sqrt{0.00207}\approx0.04550\). The standardized cutoff is \(z=(0.47357-0.46)/0.04550\approx0.2982\). Therefore, \(\beta(0.46)\approx P(Z<0.2982)\approx0.6172\), and \(\text{Power}(0.46)\approx1-0.6172=0.3828\).
The same test has estimated power about \(0.8456\) at \(p=0.52\), but only about \(0.3828\) at \(p=0.46\). The smaller departure from the null is harder for the test to distinguish from ordinary sample-to-sample variation. In context, if the true bus-use proportion is 46%, the procedure would reject the 40% null in about 38.28% of repeated samples under the stated assumptions.
Worked Example: A Smaller Alpha Changes Power
Suppose the city analyst keeps \(H_0:p=0.40\), \(H_a:p>0.40\), and \(n=120\), but uses \(\alpha=0.01\). Estimate power when the true proportion is \(p_1=0.52\). The same random-sampling and population-size assumptions apply. The null Large Counts values remain 48 and 72, and the alternative values remain 62.4 and 57.6, so both models meet the Large Counts condition. The 10% condition is also unchanged.
For an upper-tail test with \(\alpha=0.01\), use \(z^*=2.326\). The null standard deviation is still \(\sqrt{0.40(0.60)/120}\approx0.04472\). The rejection cutoff is \(0.40+2.326(0.04472)\approx0.50402\). This is higher than the cutoff of \(0.47357\) for \(\alpha=0.05\), so a sample proportion must be more extreme to reject \(H_0\).
At \(p_1=0.52\), the alternative standard deviation remains approximately \(0.04561\). The standardized cutoff is \(z=(0.50402-0.52)/0.04561\approx-0.3503\). Thus \(\beta(0.52)\approx P(Z<-0.3503)\approx0.3630\), so \(\text{Power}(0.52)\approx1-0.3630=0.6370\).
At the same true proportion and sample size, estimated power is about \(0.8456\) for \(\alpha=0.05\) and \(0.6370\) for \(\alpha=0.01\). The smaller significance level makes rejection harder, reducing the chance to detect this specified increase. The corresponding chance of a Type II error increases. This illustrates the tradeoff discussed in “Choosing Alpha Based on Error Consequences.”
What Power Does—and Does Not—Say
Power describes a planned test’s behavior over repeated samples under a specified alternative. It is not the probability that the alternative hypothesis is true, the probability that a particular result is correct, or the probability that the test will detect every possible departure from the null. Always state the true value at which power is being considered.
Power also does not establish cause and effect by itself. In the bus-use example, the test concerns whether a population proportion exceeds a benchmark. A random sample can support an inference about that population proportion, but it does not show that a particular policy caused bus use to increase. A causal question requires an appropriate randomized comparison.
A test’s power depends on its full setup, including the sample size, significance level, and specified true parameter value. When comparing powers, say what is being held fixed. For example, the second example changed the alternative value while keeping the test plan fixed; the third changed alpha while holding the true proportion and sample size fixed.
Common Mistakes and AP Exam Tip
- Leaving out the specified truth: A statement such as “the power is 0.80” is incomplete when \(H_a\) contains many possible values. State the value of \(p_1\) and the test setup.
- Reversing power and beta: Beta is the probability of failing to reject a false null; power is the probability of rejecting it. Show the relationship \(1-\beta(p_1)\) to make the complement clear.
- Describing power as a probability that a hypothesis is true: Power is conditional on a specified true parameter value. A full-credit interpretation says, “If the true proportion is \(p_1\), this test will reject \(H_0\) in about this fraction of repeated samples.”
- Using the null distribution to calculate power directly: The null distribution sets the rejection cutoff. To estimate power, use the sampling distribution at the specified alternative to find the probability of rejecting.
- Claiming causation from a sample alone: A test comparing a proportion with a benchmark does not show that a program or policy caused a change. Match any causal conclusion to an appropriate randomized design.
- Forgetting approximation language: When power is estimated with a normal approximation, call it approximate and check the randomization or random-sampling assumptions, the 10% condition when relevant, and Large Counts under the null and specified alternative.
Check Your Understanding
Answer each question using the meaning of power for a significance test.
- If \(\beta(0.48)=0.30\), what is the power at \(p_1=0.48\), and what does it mean in context?
- Why must a power value for a test with \(H_a:p>0.40\) be linked to a specific true proportion?
- A test has power 0.75 when the true proportion is 0.52. Does this mean there is a 75% probability that \(p=0.52\)? Explain.
- For an upper-tail test, how can you estimate power from the rejection cutoff and the sampling distribution under a specified alternative?
- If alpha is made smaller while the sample size and alternative proportion stay fixed, what generally happens to the rejection cutoff and power? Explain.