Power Depends on the Signal and the Noise
In “How Sample Size Affects Power in a t Test” and “How Significance Level Affects Power and Errors,” we held some features of a study fixed and considered how changing another feature affects power. Now we focus on two features of the population: how far the true mean is from the null value, and how much individual measurements vary.
A test is more likely to detect a real difference when the difference is large relative to the variability in the data. In this comparison, the true difference is the signal; variability is one source of noise that can make the sample mean fluctuate from sample to sample. The comparison only isolates these effects when other features—such as sample size, significance level, and test direction—are held fixed.
The raw difference \(\Delta\) keeps the response’s original units, which are essential for judging practical importance. The standardized difference helps compare how large that difference is relative to the natural spread of the measurements. Neither quantity alone determines power: power also depends on sample size, \(\alpha\), and the direction of the alternative.
For a fixed sample size, a larger true difference in the direction of the alternative generally makes the sample statistic more likely to enter the rejection region. Less population variability generally has a similar effect: it makes the sample mean less variable, so a given true difference stands out more clearly. As in “What Power of a Test Means,” power is the probability of rejecting \(H_0\) for a specified true mean that makes \(H_0\) false.
In a real t test, \(\sigma\) is generally unknown and the test uses the sample standard deviation \(s\). For planning examples, a pilot estimate or other planning value for \(\sigma\) can be used to approximate the standard error. One AP-level approximation locates the rejection boundary using a t critical value, then estimates the chance of crossing that boundary using a normal model for the sample mean. Treat resulting power values as approximate, not exact.
Reading the Relationship Between Effect and Variability
For an upper-tailed test, \(H_0:\mu=\mu_0\) versus \(H_a:\mu>\mu_0\), the rejection boundary for the sample mean is approximately \(\mu_0+t^*SE_{\text{planning}}\). Under a specified true mean, power is approximately the probability that \(\bar{x}\) exceeds that boundary. A larger true mean moves the sampling distribution farther above the boundary; a smaller standard error makes that distribution narrower.
This expression also shows why a difference in the wrong direction does not help an upper-tailed test. A true mean below \(\mu_0\) is still a false null hypothesis, but the sample mean is then less likely to exceed the upper rejection boundary. For a two-sided test, differences farther from the null in either direction can make rejection more likely; the examples here use upper-tailed tests to keep the calculations clear.
Worked Example: A Larger True Difference
Worked Example: Detecting a Battery-Life Improvement
A fictional electronics team plans to test whether a new battery-setting procedure increases average device runtime. For each of \(n=64\) devices, the team will calculate the change in runtime from the old setting to the new one. Let \(\mu\) be the population mean change, in hours. The team plans an upper-tailed one-sample t test of \(H_0:\mu=0\) versus \(H_a:\mu>0\), at \(\alpha=0.05\). A pilot study suggests a planning standard deviation of 8 hours. Compare approximate power if the true mean increase is 1, 2, or 3 hours.
State. The parameter is the population mean change in device runtime. We compare power for three specified true mean increases, holding the sample size, planning variability, significance level, and test direction fixed.
Plan. The paired changes are analyzed as one sample of differences. The differences should come from an appropriate random sample or randomized experiment; distinct device-pair observations should be independent, with the 10% condition met if sampling without replacement. The distribution of differences should not show severe skewness or outliers; with \(n=64\), the t procedure is generally robust to moderate departures from normality. For the power approximation, use the one-sided t critical value with \(df=63\), then a normal model for the sample mean.
Do. The planning standard error is:
The one-sided critical value for \(\alpha=0.05\) and \(df=63\) is approximately \(t^*=1.6694\). Since the null value is 0, the rejection boundary for \(\bar{x}\) is \(0+1.6694(1)=1.6694\) hours. Under each assumed true mean, standardize the distance from that mean to the boundary:
- For a true increase of 1 hour, \(z=(1.6694-1)/1=0.6694\). Approximate power is \(P(Z>0.6694)\approx0.2516\).
- For a true increase of 2 hours, \(z=(1.6694-2)/1=-0.3306\). Approximate power is \(P(Z>-0.3306)\approx0.6295\).
- For a true increase of 3 hours, \(z=(1.6694-3)/1=-1.3306\). Approximate power is \(P(Z>-1.3306)\approx0.9083\).
Conclude. For this planned test, approximate power increases from 0.2516 to 0.6295 to 0.9083 as the assumed true mean increase grows from 1 to 3 hours. In repeated studies under a true increase of 3 hours, for example, the test would reject the null claim of no mean increase about 90.83% of the time under this approximation. These are planning estimates, not predictions of what any one sample will show.
Worked Example: Less Variability
Worked Example: A More Consistent Water-Filter Measurement
A fictional lab plans to test whether a filter reduces the average amount of a residue in water. For each of \(n=64\) matched water samples, let the measured difference be the amount before filtering minus the amount after filtering, in milligrams per liter. The lab will test \(H_0:\mu=0\) versus \(H_a:\mu>0\) at \(\alpha=0.05\). It expects a true mean reduction of 2 milligrams per liter. Compare approximate power with planning standard deviations of 8 and 4 milligrams per liter.
State. The parameter is the population mean reduction in residue. The true reduction, sample size, significance level, and test direction stay fixed; only the planning standard deviation changes.
Plan. Analyze the matched before-and-after measurements using one difference per sample. The matched samples should be appropriately randomly selected or assigned, and the differences from distinct samples should be independent; check the 10% condition when sampling without replacement. Assess the differences for severe skewness or outliers. With \(n=64\), use the same t critical value, \(1.6694\) for \(df=63\). Then approximate the sample mean’s distribution using each planning standard deviation.
Do. If the planning standard deviation is 8, the standard error is \(8/\sqrt{64}=1\) milligram per liter. The rejection boundary for \(\bar{x}\) is \(0+1.6694(1)=1.6694\). Thus, \(z=(1.6694-2)/1=-0.3306\), and approximate power is \(P(Z>-0.3306)\approx0.6295\).
If the planning standard deviation is 4, the standard error is \(4/\sqrt{64}=0.5\) milligram per liter. The boundary is \(0+1.6694(0.5)=0.8347\). Thus, \(z=(0.8347-2)/0.5=-2.3306\), and approximate power is \(P(Z>-2.3306)\approx0.9901\).
Conclude. With the same 2-milligram-per-liter true mean reduction and the same design, reducing the planning standard deviation from 8 to 4 raises approximate power from 0.6295 to 0.9901. The smaller variability makes the standard error smaller, so the expected sample mean is farther above the rejection boundary in standard-error units.
Worked Example: The Direction of the Difference Matters
Worked Example: A Sensor Test for an Increase
A fictional manufacturer plans an upper-tailed one-sample t test to determine whether a sensor’s average reading has increased above a target of 50 units. The plan uses \(n=64\), \(\alpha=0.05\), and a planning standard deviation of 8 units. Suppose the true population mean is 49 units. Estimate the power for this specified alternative.
State. Here \(H_0:\mu=50\), \(H_a:\mu>50\), and the specified true difference is \(\Delta=49-50=-1\) unit. The null hypothesis is false, but the true mean is below the target rather than above it.
Plan. Assume an appropriate random sample or randomized experiment, independent observations, and the 10% condition if sampling without replacement. The sample should not show severe skewness or outliers; the sample size is 64. Use \(t^*=1.6694\) with \(df=63\), and approximate the sample mean’s distribution with a normal model centered at 49 units.
Do. The planning standard error is \(8/\sqrt{64}=1\) unit. The upper-tail rejection boundary is \(50+1.6694(1)=51.6694\) units. Standardizing this boundary under the true mean gives \(z=(51.6694-49)/1=2.6694\). Therefore, approximate power is \(P(Z>2.6694)\approx0.0038\).
Conclude. For a true mean of 49 units, the approximate probability of rejection is 0.0038. It is very unlikely to reject \(H_0\), because the true mean is in the direction opposite to the alternative. Conventional power for this upper-tailed test refers to true means above 50; the result illustrates why a probability statement must specify the true mean and the test direction.
Common Mistakes and AP Exam Tips
- Saying any larger difference increases power: For a one-sided test, the difference must grow in the direction of the alternative. A true mean on the opposite side of the null can produce very low power, as in the sensor example.
- Comparing raw differences without considering spread: A difference of 2 units may be large relative to a standard deviation of 4, but modest relative to a standard deviation of 8. Use the original units to describe the effect and the standardized difference to compare it with variability.
- Changing several features at once: If sample size, \(\alpha\), true difference, and variability all change, you cannot attribute a power difference to just one of them. State what is held fixed.
- Treating approximate planning power as exact: The examples use a planning standard deviation and a normal approximation after locating the boundary with a t critical value. Identify the result as approximate and specify the assumed true mean.
- Confusing statistical detection with practical importance: High power means a test is likely to reject a false null for the specified alternative. It does not establish that the effect is large enough to matter in practice. As discussed in “Is a Statistically Significant Difference in Means Practically Important?” and “Effect Size in Terms of the Original Units,” interpret the raw difference in context.
A strong AP response explains the comparison: “With sample size, significance level, and test direction held fixed, a larger true mean increase raises power,” or “For the same true mean difference, a smaller population standard deviation generally raises power by reducing the standard error.” If giving a numerical power estimate, name the assumed true mean, explain the planning approximation, and interpret the probability in repeated studies.
Check Your Understanding
Assume sample size, significance level, and test direction stay fixed unless a question says otherwise.
- For an upper-tailed test, what happens to power if the true mean moves farther above the null value while variability stays the same?
- A test has a true mean difference of 3 units. Explain how reducing the population standard deviation from 12 units to 6 units would generally affect power.
- What does the standardized difference \(\Delta/\sigma\) describe, and why should the raw difference still be reported in context?
- Why can an upper-tailed test have low power when the true mean differs from the null in the lower direction?
- When a power estimate uses a planning standard deviation and a normal approximation, how should the result be described?