Tutorials › AP Statistics › How Sample Size and Spread Affect the Test Result

Two-sample t hypothesis tests · Tutorial 735 of 1000

How Sample Size and Spread Affect the Test Result

Trace how changing either group’s sample size or standard deviation changes the evidence in a two-sample t test.

Intermediate 9 min read

What You'll Learn

  • Explain how increasing sample size affects the standard error and the magnitude of the t statistic.
  • Explain why greater sample standard deviations tend to weaken evidence for a fixed difference in sample means.
  • Use each group’s separate variance contribution to compare two-sample t test results.
  • Connect changes in the t statistic and degrees of freedom to changes in the p-value.
  • Describe what a changed p-value does—and does not—say about the null hypothesis.

Same Difference, Different Evidence

A two-sample t test compares an observed difference in sample means with the amount of sampling variation we would expect if the population means were equal. As covered in “The Two-Sample t Test Statistic,” that comparison is the test statistic \(t\). A fixed difference in sample means does not guarantee a fixed test result: sample sizes and sample standard deviations affect the estimated standard error, and the sample sizes also affect the degrees of freedom.

This tutorial focuses on how those ingredients change the t statistic and the p-value. The calculations use the unpooled two-sample t procedure, which keeps the two groups’ variability contributions separate. The group order and observed difference stay fixed when comparing results, so any change in the test result can be traced to the sample sizes or spreads.

Formula: For a test of \(H_0:\mu_1-\mu_2=0\), the unpooled two-sample t statistic is \[ t=\frac{\bar{x}_1-\bar{x}_2}{\sqrt{s_1^2/n_1+s_2^2/n_2}}. \] The denominator is the estimated standard error. The p-value also depends on the alternative hypothesis and the degrees of freedom used for the t distribution.

How Sample Size Changes the Test

The standard error contains one variance contribution from each group: \(s_1^2/n_1\) and \(s_2^2/n_2\). If a group’s sample size increases while its sample standard deviation stays the same, that group’s contribution gets smaller. The total standard error therefore decreases, and a fixed observed difference becomes more standard errors away from zero. In other words, the absolute value of \(t\) increases.

A larger absolute t statistic generally corresponds to a smaller p-value. There is one detail to keep in mind: sample size also changes the degrees of freedom, so the precise p-value depends on both \(t\) and the relevant t distribution. The overall pattern is still useful: for the same observed difference and similar spreads, larger samples usually provide stronger evidence against the null hypothesis.

Key idea: With the observed difference and sample spreads held fixed, increasing a sample size reduces its contribution to the standard error. This generally increases \(|t|\) and decreases the p-value, though the exact p-value uses the degrees of freedom as well.

How Sample Spread Changes the Test

The sample standard deviations \(s_1\) and \(s_2\) measure the spreads of the observed values in the two groups. Because they are squared in the standard error formula, a larger standard deviation increases that group’s variance contribution. The standard error rises, and the same observed difference becomes fewer standard errors from zero. Thus \(|t|\) usually decreases and the p-value usually increases.

The two contributions matter separately. A change in \(s_1\) affects \(s_1^2/n_1\), not \(s_2^2/n_2\); a change in \(s_2\) affects the other term. The effect of a change in spread also depends on sample size. A variance contribution is divided by its group’s sample size, so the same increase in spread has a smaller impact when that group has a larger \(n\).

The sign of \(t\) comes from the order of subtraction, \(\bar{x}_1-\bar{x}_2\). Changing a standard deviation or sample size affects the magnitude of the statistic, not the sign, as long as the observed means and their order do not change. For a two-sided test, the p-value reflects how far \(|t|\) is from zero. For a one-sided test, the direction of \(t\) matters too, because the alternative hypothesis determines which tail is used.

Worked Examples

Worked Example: Larger Samples, Smaller P-Value

In two hypothetical randomized experiments, researchers compare the mean time two designs of a rechargeable garden sensor operate, in hours. In each experiment, Design 1’s sample mean is 46 hours and Design 2’s is 40 hours, so the observed difference is 6 hours. The sample standard deviations are 6 hours in both groups. The first experiment assigns 5 sensors to each design; the second assigns 20 sensors to each design. Do the test results change when the sample sizes increase?

State: For each experiment, let \(\mu_1\) and \(\mu_2\) be the true mean operating times for sensors using Designs 1 and 2, respectively. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\).

Plan: Use an unpooled two-sample t test. In each hypothetical experiment, sensors are randomly assigned to one design, and each sensor receives only one design, so the groups are independent. With 5 sensors per group in the smaller experiment, the response distributions need close inspection; assume plots show no strong skewness or outliers. The same distribution check is satisfactory for the 20-per-group experiment. Because these are randomized treatment groups rather than samples drawn without replacement from a finite population, a 10% condition is not needed.

Do: In both experiments, the difference in sample means is \(46-40=6\) hours. With 5 sensors in each group, the standard error and t statistic are

$$ SE=\sqrt{\frac{6^2}{5}+\frac{6^2}{5}} =\sqrt{7.2+7.2} =\sqrt{14.4} \approx3.7947\text{ hours}, \qquad t=\frac{6}{3.7947}\approx1.581 $$

The Welch degrees of freedom are 8. The two-sided p-value is approximately \(0.1528\), rounded to four decimal places. For 20 sensors in each group, the standard error and test statistic are

$$ SE=\sqrt{\frac{6^2}{20}+\frac{6^2}{20}} =\sqrt{1.8+1.8} =\sqrt{3.6} \approx1.8974\text{ hours}, \qquad t=\frac{6}{1.8974}\approx3.162 $$

The Welch degrees of freedom are 38, and the two-sided p-value is approximately \(0.0030\), rounded to four decimal places. These p-values can be obtained with a calculator’s 2-SampTTest function using the unpooled setting and the two-sided alternative.

Conclude: The larger samples make the standard error smaller, so the same 6-hour observed difference is farther from zero in standard-error units. The p-value falls from about 0.1528 to about 0.0030. At \(\alpha=0.05\), the smaller experiment fails to reject \(H_0\), while the larger experiment rejects \(H_0\). The larger experiment provides convincing evidence of a difference in mean operating times for the experimental units; the smaller one does not provide convincing evidence of a difference. Neither result proves that the population means are equal or different.

Worked Example: More Spread, Larger P-Value

Consider two hypothetical randomized comparisons of the same two types of reusable water bottles. The response is the volume of water, in milliliters, that a bottle keeps cool during a fixed test. In each comparison, the sample mean is 6 milliliters higher for Type 1 than for Type 2, and each group has 16 bottles. In one comparison, both sample standard deviations are 4 milliliters; in the other, both are 8 milliliters. How does the larger spread affect the two-sided test?

State: Let \(\mu_1\) and \(\mu_2\) be the true mean cooling volumes for Types 1 and 2. For each comparison, test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\). The observed difference is 6 milliliters in each.

Plan: Use an unpooled two-sample t test. Each bottle is assigned to one type for the hypothetical test, so the groups are independent. With 16 bottles per group, assume plots of the response in both groups show no severe skewness or outliers. These conditions support the t procedure. No 10% condition is needed for this randomized assignment design.

Do: When both standard deviations are 4 milliliters,

$$ SE=\sqrt{\frac{4^2}{16}+\frac{4^2}{16}} =\sqrt{1+1} =\sqrt{2} \approx1.4142\text{ milliliters}, \qquad t=\frac{6}{1.4142}\approx4.243 $$

The Welch degrees of freedom are 30, and the two-sided p-value is approximately \(0.0002\). When both standard deviations are 8 milliliters,

$$ SE=\sqrt{\frac{8^2}{16}+\frac{8^2}{16}} =\sqrt{4+4} =\sqrt{8} \approx2.8284\text{ milliliters}, \qquad t=\frac{6}{2.8284}\approx2.121 $$

The Welch degrees of freedom are again 30, and the two-sided p-value is approximately \(0.0422\), rounded to four decimal places.

Conclude: The larger spreads increase the standard error from about 1.4142 to 2.8284 milliliters. That cuts the magnitude of \(t\) and increases the p-value, even though the observed difference and sample sizes are unchanged. Both p-values are below 0.05, so both tests reject \(H_0\) at that level, but the evidence against \(H_0\) is weaker in the comparison with more spread.

Worked Example: One Group’s Spread Changes

A hypothetical randomized study compares two settings for a small water pump. The response is the volume pumped in one minute, in liters. Both samples have 16 pumps, and the observed difference in sample means is 6 liters, with the mean for Setting 1 higher. The standard deviation for Setting 1 is 4 liters in both comparisons. Setting 2’s standard deviation is 4 liters in the first comparison and 8 liters in the second. Calculate how changing only \(s_2\) affects the test.

State: Let \(\mu_1\) and \(\mu_2\) be the true mean volumes pumped per minute under Settings 1 and 2. In each comparison, test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\).

Plan: Use an unpooled two-sample t test. The hypothetical study randomly assigns pumps to settings, and each pump receives only one setting, so the groups are independent. With 16 pumps per group, assume plots show no severe skewness or outliers in either group. The randomized design does not require a 10% condition.

Do: When \(s_1=4\) and \(s_2=4\), the standard error is \(\sqrt{1+1}=\sqrt{2}\approx1.4142\) liters and \(t=6/1.4142\approx4.243\). When only \(s_2\) changes to 8, its variance contribution changes from \(4^2/16=1\) to \(8^2/16=4\), while Setting 1’s contribution remains \(4^2/16=1\). Thus,

$$ SE=\sqrt{\frac{4^2}{16}+\frac{8^2}{16}} =\sqrt{1+4} =\sqrt{5} \approx2.2361\text{ liters}, \qquad t=\frac{6}{2.2361}\approx2.683 $$

The Welch degrees of freedom are approximately 22.1. The two-sided p-value is approximately \(0.0136\), rounded to four decimal places. When both spreads are 4, the degrees of freedom are 30 and the p-value is approximately \(0.0002\).

Conclude: Increasing only Setting 2’s spread raises the total standard error, decreases \(|t|\), and raises the p-value. Both results still reject \(H_0\) at \(\alpha=0.05\), but the result with greater spread provides less evidence against equal population means.

Common Mistakes and AP Exam Tips

A good comparison explains the chain from the data to the test result: sample sizes and spreads determine the standard error; the standard error and observed difference determine \(t\); and \(t\), the alternative, and the degrees of freedom determine the p-value.

  • Saying that a larger sample changes the observed difference: It might change the difference in a new sample, but in a controlled comparison like these examples, the difference is held fixed to isolate the effect of sample size.
  • Claiming that larger \(s\) makes the t statistic larger: For a fixed observed difference, larger spread increases the standard error and generally reduces \(|t|\). The sign is still determined by the order of the sample means.
  • Combining the spreads before calculating: Use \(\sqrt{s_1^2/n_1+s_2^2/n_2}\). Each group contributes its own variance divided by its own sample size.
  • Claiming that the p-value depends only on \(t\): A p-value also uses the t distribution determined by the degrees of freedom, as well as the direction of the alternative hypothesis.
  • Calling the p-value the probability that the null hypothesis is true: It is calculated assuming the null hypothesis is true and describes the chance of a test statistic at least as extreme as the observed one.
  • Equating a larger p-value with proof of no difference: A larger p-value means the data provide less evidence against the null under the test. It does not establish that the population means are equal.

For full credit, show the standard error calculation, the signed t statistic, and the p-value using the correct alternative and degrees of freedom. When comparing scenarios, state what is being held constant and explain the direction of change. A p-value can change without changing the decision at a particular significance level; for example, two different p-values may both be below 0.05.

Key takeaway: For a fixed observed difference in sample means, larger sample sizes generally reduce the standard error and p-value, while larger sample spreads generally increase the standard error and p-value. Each group’s contribution is \(s_i^2/n_i\), and the p-value also depends on the alternative hypothesis and degrees of freedom.

Check Your Understanding

Assume the observed difference in sample means and the alternative hypothesis stay the same unless a question says otherwise.

  1. If \(s_1\), \(s_2\), and the observed difference stay fixed while \(n_1\) increases, what happens to group 1’s variance contribution to the standard error?
  2. If the standard error increases but the observed difference stays fixed, what happens to the magnitude of \(t\)?
  3. Why can two tests with the same t statistic have slightly different p-values?
  4. In the third worked example, which variance contribution changed when Setting 2’s standard deviation increased, and by how much?
  5. A p-value decreases after sample size increases but remains above \(\alpha=0.05\). What is the appropriate decision at that level, and what does it say about the evidence?