Power Describes a Test’s Ability to Detect a False Null
In “Alpha as the Type I Error Rate in t Tests,” you learned that alpha describes how often a test rejects a true null hypothesis over repeated samples. Now consider what happens when the null hypothesis is false. A test may reject it, or it may fail to reject it. The probability of rejection in this situation is called the test’s power.
The phrase “the particular true mean being considered” is important. A null hypothesis such as \(H_0:\mu=50\) is false for many possible values of \(\mu\): it might be 49, 52, or 60. The probability of rejecting \(H_0\) can differ across those possibilities. So, when interpreting power, state which population mean is assumed to be true.
Power is a long-run probability. It describes the proportion of repeated samples that would lead the specified test to reject \(H_0\), if the population mean really were the specified false-null value and the test’s conditions and model assumptions held. It does not tell us for certain what will happen in one future sample.
The vertical bar means “given that.” For example, the probability in this definition is conditional on a particular mean actually being true. Power is not the probability that the mean has that value, and it is not the probability that \(H_0\) is false.
Interpreting Power in Context
A power statement should connect three things: the test decision, the particular true population value, and the context. Suppose a test of a mean has power 0.80 when the true population mean is 8 hours. A careful interpretation is: “If the true mean battery life is 8 hours, this test would reject the null hypothesis in about 80% of repeated samples, assuming the test conditions hold.”
That statement does not claim that the mean is 8 hours. It describes how the test would perform if 8 hours were the true mean. Nor does it say that an individual test has an 80% chance of being correct. Each particular sample produces a test decision; power describes the long-run behavior of the procedure under a specified population truth.
Power is also different from the p-value. The p-value is the probability, assuming \(H_0\) is true, of obtaining a result at least as extreme as the observed result. Power is a probability about the test’s repeated-sample decisions when a specified alternative value is true. A p-value is available after observing data; power is tied to a test design and a possible true mean.
Worked Examples: Giving Power a Meaning
Worked Example: Detecting a Change in Mean Battery Life
A fictional testing team plans a one-sample t test of \(H_0:\mu=10\) hours against \(H_a:\mu\ne10\), using \(\alpha=0.05\) and a sample of 25 batteries. Here, \(\mu\) is the true mean battery life for the population represented by the sampling process. Suppose a planning simulation estimates that this test would reject \(H_0\) in 800 of 1,000 repeated samples if the true mean battery life were 12 hours.
Identify the setting. The null value is 10 hours, while the particular false-null value under consideration is 12 hours. The test is two-sided, and the power estimate concerns this test with its stated sample size and significance level.
Interpret the estimate. The simulated proportion of rejections is:
So the estimated power is 0.80, or 80%. If the true mean battery life is 12 hours, this testing procedure would reject the claim that the mean is 10 hours in about 80% of repeated samples. In about 20% of those repetitions, it would fail to reject \(H_0\), even though the mean really is 12 hours.
Keep the scope clear. This interpretation assumes the repeated samples follow the planning model and the t test conditions are appropriate. It does not establish that the true mean is 12 hours, and it does not guarantee rejection in any one sample.
Worked Example: Power for a Paired t Test of Recovery Time
A fictional clinic is planning a study in which each participant’s recovery time is recorded before and after a program. Define \(d\) as the participant’s recovery time before the program minus the time after it, in days. A positive difference means the participant recovered sooner after the program. The planned paired t test uses \(H_0:\mu_d=0\) against \(H_a:\mu_d>0\), with 20 participants and \(\alpha=0.05\).
Assume the participant pairs are obtained through an appropriate random process, the differences from distinct participants are independent, and the distribution of the differences is suitable for a paired t procedure. Suppose a planning simulation produces 647 rejections out of 1,000 repetitions when the true mean difference is 3 days.
Calculate the estimated power. The simulated proportion of rejections is:
The estimated power is 0.647, or 64.7%. In context, if the true mean reduction in recovery time is 3 days, this planned test would reject \(H_0:\mu_d=0\) in about 64.7% of repeated studies that meet the stated design and model assumptions.
This interpretation follows the direction of the alternative: the test is looking for evidence that the mean reduction is positive. The estimated power is not the probability that the program works, nor is it the probability that the mean reduction equals 3 days. It describes the test’s chance of rejecting its null under that specific assumed mean difference.
Worked Example: Power at Two Different False-Null Values
A fictional environmental team plans a two-sided one-sample t test of \(H_0:\mu=50\) against \(H_a:\mu\ne50\), where \(\mu\) is the true mean sensor reading in units. The sample size, test conditions, and \(\alpha=0.05\) are held fixed. A planning simulation gives 310 rejections out of 1,000 repetitions when the true mean is 52 units, and 880 rejections out of 1,000 repetitions when the true mean is 58 units.
Estimate power at each value. At a true mean of 52 units:
The estimated power is 31.0%. If the true mean reading is 52 units, the test would reject \(H_0\) in about 31.0% of repeated samples.
At a true mean of 58 units:
The estimated power is 88.0%. If the true mean reading is 58 units, the same test would reject \(H_0\) in about 88.0% of repeated samples.
These are two power values for the same test, not a contradiction. The true mean of 58 is farther from the null value of 50 than is 52, so the test is more likely to distinguish that larger departure from the null, under the planning model. Power must be tied to a stated alternative value; saying only “the test has power 88%” leaves out an important part of the interpretation.
Worked Example: How Changing Alpha Can Affect Power
A fictional agriculture team considers a one-sided one-sample t test of \(H_0:\mu=40\) against \(H_a:\mu>40\), where \(\mu\) is the mean crop yield in kilograms per plot. The team compares two significance levels while keeping the sample size and other design features fixed. For a specified true mean of 43 kilograms per plot, a planning simulation estimates 540 rejections out of 1,000 repetitions at \(\alpha=0.01\), and 780 rejections out of 1,000 at \(\alpha=0.05\).
Calculate and interpret both estimates. At \(\alpha=0.01\), the estimated power is:
If the true mean yield is 43 kilograms per plot, the test at this significance level would reject \(H_0\) in about 54% of repeated samples. At \(\alpha=0.05\), the estimated power is:
With the same true mean and design otherwise unchanged, the test at \(\alpha=0.05\) would reject \(H_0\) in about 78% of repeated samples. A larger alpha allows a less extreme result to lead to rejection, which can make rejection more likely when the null is false. The trade-off is that, when the null is true, the long-run Type I error rate is larger. Alpha and power describe different error situations, as discussed in “Alpha as the Type I Error Rate in t Tests.”
Common Mistakes and AP Exam Tips
- Leaving out the true mean being considered: A power statement needs a specific false-null value. Say “if the true mean is 43 kilograms per plot,” not simply “if the null is false.”
- Calling power the probability that \(H_0\) is false: Power is conditional on a specified population mean being true. It does not assign a probability to competing hypotheses.
- Confusing power with alpha: Alpha is the probability of rejecting a true null under the null model. Power is the probability of rejecting a false null at a specified true value.
- Calling power the probability a particular study will succeed: Power is a long-run probability for a repeated-sampling procedure. It does not guarantee what one sample will show.
- Reporting a power percentage without context: A full-credit interpretation names the test decision, the population parameter’s assumed true value, and the population or response being studied.
- Treating one power value as applying to every alternative: A test can have different power for different values of the population mean. State the particular alternative value connected to the power figure.
- Calling a simulation estimate exact: A proportion from a finite number of simulated samples estimates power. For example, 647 rejections in 1,000 repetitions gives an estimated power of 0.647, not a guarantee that exactly 64.7% of future sets of 1,000 studies will reject.
For a strong AP response, use a sentence such as: “If the true population mean is [specified value], this test would reject \(H_0\) in about [power] of repeated samples, assuming the conditions for the test hold.” Fill in the value, probability, and context. This makes clear both what event is being counted and what population truth the interpretation assumes.
Check Your Understanding
For each question, identify the assumed true population value and describe the repeated-sample event represented by power.
- A two-sided t test of \(H_0:\mu=25\) has power 0.72 when the true mean is 28. Interpret this power in context if \(\mu\) measures mean weekly screen time in hours.
- Why is “the test has power 80%” incomplete if no alternative mean is specified?
- A simulation produces 430 rejections in 1,000 repetitions when a specified false-null mean is true. What is the estimated power, and what does it mean?
- In a mean test, what does alpha describe, and what does power describe? State the difference in the population truth each one assumes.
- Why does a power value not mean there is that probability the null hypothesis is false?