Tutorials › AP Statistics › Effect of Unequal Sample Sizes on Standard Error

Two-sample t confidence intervals · Tutorial 713 of 1000

Effect of Unequal Sample Sizes on Standard Error

Learn how unequal sample sizes change each group’s contribution to the standard error and how to compare allocations fairly.

Intermediate 9 min read

What You'll Learn

  • Calculate how each sample contributes to the standard error of a difference in means.
  • Compare balanced and imbalanced allocations while holding total sample size fixed.
  • Explain why imbalance matters more when the groups have different variability.
  • Compare standard errors without confusing them with the margin of error or interval width.
  • Check the conditions for a two-sample t interval in a complete example.

Why the Split Between Samples Matters

In “Pooled Versus Unpooled Two-Sample Intervals,” you saw that the standard AP two-sample t interval estimates the two groups’ variance contributions separately. Here we focus on how the sample sizes affect those contributions. Two studies can have the same total number of observations but different standard errors if those observations are divided differently between the groups.

Recall the unpooled standard error from “Standard Error for a Difference in Means.” The uncertainty contributed by each sample is \(s_i^2/n_i\), and the two contributions are added before taking the square root. Thus, a small sample size can leave a relatively large contribution—even if the other group has many observations.

$$ SE_{\bar{x}_1-\bar{x}_2} = \sqrt{\frac{s_1^2}{n_1}+\frac{s_2^2}{n_2}} $$

A useful way to compare allocations is to hold the total sample size and the sample standard deviations fixed, then change \(n_1\) and \(n_2\). This is a mathematical comparison, not a promise that real samples with different sizes will have exactly the same standard deviations. In practice, the observed \(s_1\) and \(s_2\) may change from sample to sample.

Key idea: Each group’s contribution to the standard error is its sample variance divided by its sample size. Increasing one group’s sample size reduces that group’s contribution, but it does not remove the other group’s contribution.

Balanced and Imbalanced Samples

When the two groups have the same standard deviation, the standard error is smallest when a fixed total sample size is split evenly. To see why, if both standard deviations equal \(s\), then the quantity inside the square root is \(s^2(1/n_1+1/n_2)\). For a fixed total \(n_1+n_2\), the product \(n_1n_2\) is largest when the two sample sizes are equal. Since \(1/n_1+1/n_2=(n_1+n_2)/(n_1n_2)\), the sum of the reciprocal sample sizes—and therefore the standard error—is minimized by a balanced split.

This does not mean that an imbalanced design is automatically poor. Practical limits, such as the availability of participants or the cost of measurements, can make unequal sample sizes sensible. It does mean that a large total sample size alone does not tell you how precise the difference in sample means will be. You also need to know how that total is divided and how variable each group is.

Comparison rule: To isolate the effect of sample-size imbalance, keep the total \(n_1+n_2\) and the two standard deviations fixed. Then calculate \(s_1^2/n_1\) and \(s_2^2/n_2\) for each allocation, add them, and take the square root.

Worked Examples

Worked Example: Equal Variability, Unequal Allocation

An invented study compares the mean time, in minutes, that two independent groups of volunteers take to complete a task. Suppose both groups have sample standard deviations of 6 minutes, and the total sample size is 40. Compare a balanced allocation of 20 volunteers per group with an imbalanced allocation of 10 in group 1 and 30 in group 2. This example focuses only on the standard error, not on an interval or test.

With 20 volunteers in each group, the standard error is:

$$ SE_{\text{balanced}} = \sqrt{\frac{6^2}{20}+\frac{6^2}{20}} = \sqrt{1.8+1.8} = \sqrt{3.6} \approx 1.897\text{ minutes} $$

With 10 volunteers in group 1 and 30 in group 2, it is:

$$ SE_{\text{imbalanced}} = \sqrt{\frac{6^2}{10}+\frac{6^2}{30}} = \sqrt{3.6+1.2} = \sqrt{4.8} \approx 2.191\text{ minutes} $$

The imbalanced allocation has a standard error about \(2.191-1.897=0.294\) minutes larger. Relative to the balanced standard error, that is about \(0.294/1.897\approx0.155\), or 15.5% larger. Both allocations have 40 volunteers total and the same assumed standard deviations. The difference comes from the way those volunteers are divided: the group with only 10 observations contributes more uncertainty.

The calculation also illustrates why one group’s large sample cannot entirely compensate for the other group’s small sample. Group 2’s contribution falls from \(36/20=1.8\) to \(36/30=1.2\), but group 1’s contribution rises from \(36/20=1.8\) to \(36/10=3.6\). The total inside the square root increases.

Worked Example: Giving More Observations to the More Variable Group

Now consider an invented comparison of two independent types of reusable water filters. Suppose measurements of the filters’ operating life have sample standard deviations \(s_1=8\) hours and \(s_2=4\) hours. A study can measure 40 filters total. Compare allocating 10 to group 1 and 30 to group 2 with reversing that allocation.

When the more variable group 1 has only 10 filters, the standard error is:

$$ SE_{10,30} = \sqrt{\frac{8^2}{10}+\frac{4^2}{30}} = \sqrt{6.4+0.5333} = \sqrt{6.9333} \approx 2.634\text{ hours} $$

When group 1 has 30 filters and group 2 has 10, the standard error is:

$$ SE_{30,10} = \sqrt{\frac{8^2}{30}+\frac{4^2}{10}} = \sqrt{2.1333+1.6} = \sqrt{3.7333} \approx 1.932\text{ hours} $$

The second allocation has a smaller standard error, even though each allocation still uses 40 filters. It places more observations in the group with greater variability, reducing that group’s relatively large variance contribution. For comparison, a balanced allocation of 20 filters per group would give:

$$ SE_{20,20} = \sqrt{\frac{64}{20}+\frac{16}{20}} = \sqrt{4} = 2.000\text{ hours} $$

Under these assumed standard deviations, the allocation of 30 to the more variable group and 10 to the less variable group gives a smaller standard error than either the reversed allocation or the balanced allocation. This comparison explains why equal sample sizes are not always the most precise allocation when the groups have different variability. It does not establish a universal design rule: costs, access to units, and the purpose of the study can matter too.

Worked Example: Comparing Standard Errors and a Two-Sample Interval

An invented battery-testing study compares mean operating life, in hours, for two independent battery models. A random sample of 12 batteries from model 1 has \(\bar{x}_1=52\) hours and \(s_1=4\) hours. A random sample of 30 batteries from model 2 has \(\bar{x}_2=49\) hours and \(s_2=5\) hours. We will calculate a 95% two-sample t interval and compare its standard error with a hypothetical balanced allocation of 21 batteries per group, keeping the total sample size at 42 and the standard deviations fixed for the comparison.

State: The parameter is \(\mu_1-\mu_2\), the true mean operating life for model 1 minus the true mean operating life for model 2, in hours.

Plan: Use an unpooled two-sample t interval, as in “The Form of a Two-Sample t Interval.” The model groups are independent because each battery belongs to one model and measurements are not paired. The samples are described as random. If each model’s population contains at least 120 and 300 batteries, respectively, each sample is no more than 10% of its population. The model 1 sample is small, so its plot should show no strong skewness or outliers; assume it does. The model 2 sample has 30 observations, and assume its plot also shows no extreme skewness or outliers. These checks support using the interval.

Do: The difference in sample means is \(52-49=3\) hours. The unpooled standard error is:

$$ SE = \sqrt{\frac{4^2}{12}+\frac{5^2}{30}} = \sqrt{1.3333+0.8333} = \sqrt{2.1667} \approx 1.472\text{ hours} $$

Using the Welch formula, the degrees of freedom are approximately:

$$ df = \frac{(1.3333+0.8333)^2} {\frac{1.3333^2}{11}+\frac{0.8333^2}{29}} \approx 25.3 $$

For 95% confidence and about 25.3 degrees of freedom, \(t^*\approx2.058\). The margin of error is \(2.058(1.472)\approx3.029\) hours, so the interval is:

$$ 3\pm3.029=(-0.029,\ 6.029)\text{ hours} $$

For the hypothetical balanced allocation, \(n_1=n_2=21\), the standard error with the same standard deviations would be:

$$ SE_{\text{balanced}} = \sqrt{\frac{4^2}{21}+\frac{5^2}{21}} = \sqrt{\frac{41}{21}} \approx 1.397\text{ hours} $$

That is about \(1.472-1.397=0.075\) hours lower than the standard error for the 12-and-30 allocation. This comparison isolates the effect of the sample-size split under the stated assumptions. It does not mean the two intervals would differ only by this standard-error amount: an interval’s margin of error also depends on its critical value, which depends on the degrees of freedom.

Conclude: We are 95% confident that the true mean operating life for model 1 minus the true mean operating life for model 2 is between about \(-0.029\) and \(6.029\) hours. The interval is centered at the observed difference of 3 hours, and its margin of error reflects the standard error as well as the t critical value. The comparison shows that, with the assumed standard deviations and total sample size, the balanced allocation has the smaller standard error.

Common Mistakes and AP Exam Tips

  • Comparing only the total sample sizes: A total of 40 observations does not specify how much each group contributes to the standard error. Write both \(s_1^2/n_1\) and \(s_2^2/n_2\).
  • Putting \(n_i\) outside the square root incorrectly: The unpooled formula is \(\sqrt{s_1^2/n_1+s_2^2/n_2}\). Do not add standard deviations or divide a single combined quantity by the total \(n\).
  • Assuming a bigger sample in one group makes the other contribution disappear: Each group contributes its own variance term. A very large \(n_2\) makes \(s_2^2/n_2\) smaller, but the \(s_1^2/n_1\) term remains.
  • Claiming imbalance always has the same effect: The impact depends on both sample sizes and both standard deviations. When the groups’ variability differs, it matters which group receives more observations.
  • Equating standard error with margin of error: For a t interval, the margin of error is \(t^*SE\). A smaller standard error generally helps narrow the interval, but changes in degrees of freedom can also change \(t^*\).
  • Treating a controlled comparison as observed fact: When you hold sample standard deviations fixed while changing sample sizes, say that this is a comparison under an assumption. Actual samples may have different standard deviations.

For full credit on a calculation, show the two separate variance contributions, their sum, and the square root, with units. For a full interval, also state the parameter and confidence level, check the conditions in context, use the appropriate unpooled standard error and Welch degrees of freedom, and interpret the interval in the stated subtraction order. When asked about allocation, explain which group’s contribution changes rather than saying only that “more data lowers the standard error.”

Key takeaway: Sample-size imbalance changes the two contributions \(s_1^2/n_1\) and \(s_2^2/n_2\) separately. With equal variability and a fixed total sample size, a balanced split minimizes the standard error. With unequal variability, placing more observations in the more variable group can reduce it.

Check Your Understanding

Use the unpooled standard-error formula and keep track of each group’s contribution.

  1. Two groups have the same sample standard deviation of 5 units. With 30 observations total, compare the standard errors for allocations of 15 and 15 versus 10 and 20.
  2. For an allocation with \(s_1=7\), \(n_1=10\), \(s_2=3\), and \(n_2=30\), write each group’s variance contribution before calculating the standard error.
  3. If one group’s sample size increases while its standard deviation is held fixed, what happens to that group’s contribution \(s_i^2/n_i\)?
  4. Why might a balanced allocation not give the smallest standard error when the two groups have different standard deviations?
  5. Why is a smaller standard error not, by itself, enough information to compare the widths of two t intervals?