Why Sample Size Changes Power
In “Power and Type II Error Probability,” you learned that power is the probability a test rejects \(H_0\) when a particular false-null value is true. This tutorial asks how that probability changes when the sample size changes. To make a fair comparison, hold the true difference, the variability, the significance level, and the direction of the alternative fixed.
For a t test about a mean, a larger sample generally makes the sample mean more precise. The standard error decreases as \(n\) increases, so a fixed true difference is larger relative to the standard error. The test is then more likely to reject \(H_0\). The conclusion is about the test’s long-run behavior; a larger sample does not guarantee that any particular study will reject the null hypothesis.
For a one-sample t test, the standard error is estimated by \(s/\sqrt{n}\), where \(s\) is the sample standard deviation. For paired data, use the standard deviation of the pairwise differences and the number of pairs. For an unpooled two-sample t procedure, the estimated standard error of the difference in sample means is \(\sqrt{s_1^2/n_1+s_2^2/n_2}\). These formulas show why more observations usually reduce uncertainty, provided the variability is comparable.
Power calculations require a value for the difference that is assumed to be truly present. They also require a planning estimate of variability. In an actual planning problem, a sample standard deviation from earlier or pilot data may be used as an estimate of the population standard deviation. It is not a known population value, so power found this way is an approximation.
Comparing Power with a t-Test Rejection Region
A two-sided t test at significance level \(\alpha\) rejects \(H_0\) when the test statistic falls far enough into either tail. Equivalently, for a one-sample test of \(H_0:\mu=\mu_0\), the sample mean must be sufficiently far from \(\mu_0\). A larger sample reduces the standard error, which usually moves the rejection boundaries closer to \(\mu_0\) when expressed in the original units.
For AP Statistics planning comparisons, a useful approximation is to find those boundaries using a t critical value and then approximate the sample mean’s distribution under the specified true mean with a normal distribution. This approximation is not an exact t-test power calculation: a t procedure estimates variability from the data, and the distribution of its test statistic under a false null is not exactly standard normal. The examples identify the approximation and use the planning standard deviation in place of the unknown population standard deviation.
A useful connection is to confidence intervals. For a two-sided t test at level \(\alpha\), the matching \(100(1-\alpha)\%\) confidence interval excludes the null value exactly when the test rejects \(H_0\). Larger samples generally produce narrower intervals, making it more likely that an interval will exclude the null value when the specified true difference is nonzero. This connection is a way to understand the same change in test behavior, not a different definition of power.
Worked Examples: One-Sample t Tests
Worked Example: Comparing Two Sample Sizes for a Mean Difference
A fictional recreation center is studying the mean change in the number of minutes people spend walking after a new route is introduced. Define each person’s change as “minutes after minus minutes before,” so positive values represent increases. A pilot sample has \(\bar{x}=4.0\) minutes and \(s=8.0\) minutes. For planning, suppose the true mean change is 4 minutes and use 8 minutes as the estimate of the population standard deviation. Compare the approximate power of two-sided one-sample t tests of \(H_0:\mu=0\) versus \(H_a:\mu\ne0\), both at \(\alpha=0.05\), for \(n=16\) and \(n=64\).
State. The specified true mean change is 4 minutes, the same for both designs. The pilot value \(s=8.0\) minutes is used as the planning estimate of variability; it is not asserted to be the known population standard deviation.
Plan. Each design uses a one-sample t test of the mean change. For each sample size, calculate the standard error and the two-sided t critical value. Approximate power by finding the probability, under a normal model centered at 4 minutes with standard deviation equal to the planning standard error, that the sample mean is beyond either rejection boundary. Assume the differences are obtained using an appropriate random process, are independent (including the 10% condition when sampling without replacement), and have a distribution for which t inference is reasonable. With \(n=16\), inspect the differences for strong skewness or outliers; with \(n=64\), the t procedure is generally more robust to moderate departures from normality, though severe outliers remain a concern.
Do for \(n=16\). The estimated standard error is:
With \(df=15\), the two-sided 0.05 critical value is approximately \(t^*=2.131\). Thus, the test rejects when the sample mean change is below or above these boundaries:
Under the planning approximation, the sample mean has a normal distribution centered at 4 minutes with standard deviation 2 minutes. The upper-tail probability is \(P(Z>(4.262-4)/2)=P(Z>0.131)\approx0.4479\). The lower-tail probability is \(P(Z<(-4.262-4)/2)=P(Z<-4.131)\approx0.0000\). Adding the two tails gives approximate power \(0.4479+0.0000=0.4479\), rounded to four decimal places.
Do for \(n=64\). The estimated standard error is:
With \(df=63\), the two-sided 0.05 critical value is approximately \(t^*=1.998\). The rejection boundaries are \(0\pm1.998(1)\), or \(-1.998\) and \(1.998\) minutes. Under the planning approximation, the sample mean is normal with mean 4 and standard deviation 1. The upper-tail probability is \(P(Z>(1.998-4)/1)=P(Z>-2.002)\approx0.9774\). The lower-tail probability is \(P(Z<(-1.998-4)/1)=P(Z<-5.998)\approx0.0000\). Approximate power is \(0.9774+0.0000=0.9774\).
Conclude. Under the stated planning assumptions, increasing the sample size from 16 to 64 raises approximate power from 0.4479 to 0.9774 for detecting a true mean increase of 4 minutes. In repeated samples of 64 people, the test would reject the null claim of no mean change much more often than it would with 16 people. These are approximate planning probabilities, not guarantees for either study.
A Second Comparison: A One-Sided Test
Worked Example: Detecting an Increase in a Quantitative Measurement
A fictional greenhouse team tests whether a new lighting schedule increases the mean daily growth of seedlings. A pilot sample has \(\bar{x}=3.2\) millimeters of growth above a comparison benchmark and \(s=6.0\) millimeters. For planning, use 6.0 millimeters as the estimated population standard deviation and suppose the true mean increase is 3 millimeters. Compare one-sided one-sample t tests of \(H_0:\mu=0\) versus \(H_a:\mu>0\), at \(\alpha=0.05\), using \(n=25\) and \(n=100\).
For \(n=25\), the estimated standard error is \(6/\sqrt{25}=1.2\) millimeters. With \(df=24\), the one-sided t critical value is approximately 1.711. The rejection boundary for the sample mean is therefore:
Under the planning approximation, the sample mean is normal with mean 3 and standard deviation 1.2. The approximate power is the probability of exceeding 2.0532:
For \(n=100\), the estimated standard error is \(6/\sqrt{100}=0.6\) millimeters. With \(df=99\), the one-sided critical value is approximately 1.660, giving a rejection boundary of \(0+1.660(0.6)=0.996\) millimeters. The approximate power is:
With the assumed true increase and variability held fixed, the larger sample makes the test much more likely to reject \(H_0\). Approximate power rises from 0.7851 to 0.9996. This result applies to the specified true increase of 3 millimeters and to the one-sided test shown; it is not one universal power value for every possible increase.
Two-Sample t Tests: More Units in Each Group
For a two-sample t test, increasing the sample size in one or both independent groups generally reduces the standard error of the difference in sample means. The comparison should say whether the stated sample size means the total across both groups or the number in each group. In the example below, \(n\) is the number of units in each group.
Worked Example: Comparing Two Treatment Groups
A fictional community garden compares the mean weekly mass of tomatoes from plants receiving two different watering schedules. The pilot summaries are \(\bar{x}_A=38\) ounces and \(s_A=10\) ounces for schedule A, and \(\bar{x}_B=32\) ounces and \(s_B=10\) ounces for schedule B. For planning, assume the true difference \(\mu_A-\mu_B\) is 6 ounces and each group’s population standard deviation is about 10 ounces. Compare one-sided unpooled two-sample t tests of \(H_0:\mu_A-\mu_B=0\) versus \(H_a:\mu_A-\mu_B>0\), at \(\alpha=0.05\), with 25 and 100 plants in each group.
The pilot mean difference is \(38-32=6\) ounces. As in “One-Sample t Versus Two-Sample t,” the groups are independent, so the parameter is the difference between two population means. Assume plants were randomly assigned to the watering schedules, each plant contributes one response, observations are independent within groups, and the group distributions have no severe skewness or outliers. Use the pilot standard deviations as planning estimates. With equal sample sizes and equal planning standard deviations, the Welch degrees of freedom are approximately \(2n-2\).
With 25 plants per group, the estimated standard error of the difference is:
The degrees of freedom are approximately 48, so the one-sided 0.05 critical value is about 1.677. The rejection boundary for \(\bar{x}_A-\bar{x}_B\) is \(0+1.677(2.828)\approx4.743\) ounces. Under the planning approximation, the sample mean difference is normal with mean 6 and standard deviation 2.828. Therefore:
With 100 plants per group, the standard error is:
The degrees of freedom are approximately 198, with one-sided critical value about 1.653. The rejection boundary is \(0+1.653(1.414)\approx2.337\) ounces. Approximate power is:
If the true difference is 6 ounces and the planning variability is appropriate, the design with 100 plants in each group has much higher approximate power than the design with 25 plants in each group. The larger design requires more plants overall, so practical decisions should also consider time, cost, and whether an effect of this size is important in context.
What a Larger Sample Does—and Does Not—Mean
The examples held the assumed true effect and variability fixed. That is essential: increasing sample size is not the only way to change power. Power also depends on the size of the true difference, the population variability, the significance level, whether the test is one-sided or two-sided, and the study design. A difference that is larger relative to the variability is generally easier to detect.
Sample size affects power through precision, not by making a true effect larger. For a one-sample mean, multiplying \(n\) by 4 cuts \(s/\sqrt{n}\) in half if \(s\) stays the same. For two independent groups, increasing the sample size in each group likewise reduces the estimated standard error. A narrower confidence interval is more likely to exclude the null value when the specified true difference is nonzero.
Power is not calculated from the observed sample mean alone. In a completed study, \(\bar{x}\) and \(s\) describe that particular sample and are used in the t test or confidence interval. A power comparison instead asks how a planned procedure would behave over many repeated samples under an assumed true population value. The pilot summaries in the examples supplied plausible planning estimates; they did not prove what the population mean or standard deviation is.
Common Mistakes and AP Exam Tips
- Changing more than sample size: To isolate the effect of \(n\), keep the assumed true difference, variability, significance level, and test direction fixed. If these change too, the power comparison cannot be attributed to sample size alone.
- Calling a larger sample a guarantee: Greater power means a higher long-run probability of rejection under the specified false-null value. It does not ensure rejection in one particular sample.
- Confusing sample size with a larger effect: More observations reduce standard error; they do not change the assumed true mean difference. State the effect in the original units and keep it fixed in the comparison.
- Forgetting which sample size is being reported: In a two-sample study, say “25 in each group” or “50 total.” The standard error depends on both group sizes.
- Treating an approximate calculation as exact: If a planning standard deviation is substituted into a normal approximation, call the result approximate and identify the assumed true mean or difference. Exact t-test power calculations account for the test statistic’s distribution under the alternative.
- Skipping conditions: A large \(n\) does not repair severe bias, dependence, or poor sampling. Describe the random process and check independence and the relevant data-shape conditions, as in “Matching Procedures to Conditions.”
For a strong AP response, identify the population parameter and test, name the fixed true difference and planning variability, and explain how the larger sample reduces the standard error and changes the long-run chance of rejecting \(H_0\). If you report power, interpret it for that specified population truth and in the context of the response’s units.
Check Your Understanding
For each question, keep the assumed true difference, the design, and the meaning of power in view.
- A one-sample t test uses the same significance level and planning standard deviation in two designs, but one design has four times as many observations. What happens to the estimated standard error, and what is the usual effect on power?
- A paired t test is planned to detect a true mean change of 5 minutes. What sample standard deviation should be used in the planning standard error: the standard deviation of all before-and-after measurements, or the standard deviation of the pairwise differences?
- A two-sample study plans 40 participants total, split equally between two groups. State the sample size in each group and explain why that detail matters for the standard error.
- Why does a power estimate based on a pilot sample standard deviation need to be described as approximate?
- A test has high power for a specified true difference. Does that mean the test is guaranteed to reject \(H_0\) in the planned study? Explain.