Tutorials › AP Statistics › Effect of Sample Size on a One-Sample t Test

One-sample t hypothesis tests · Tutorial 677 of 1000

Effect of Sample Size on a One-Sample t Test

Compare one-sample t tests with the same observed mean difference to see how sample size affects the standard error, t statistic, degrees of freedom, and p-value.

Intermediate 9 min read

What You'll Learn

  • Calculate how the standard error changes when sample size increases from 10 to 100.
  • Compare t statistics and two-sided p-values when the observed mean difference and sample standard deviation stay the same.
  • Explain how sample size affects both the test statistic and the degrees of freedom.
  • Use a complete one-sample t test to compare evidence with a stated significance level.
  • Distinguish statistical significance from the size and practical importance of a mean difference.
  • Explain why sample size alone does not determine a test result when other sample summaries change.

Same Difference, Different Evidence

A sample mean can be a fixed distance from a null mean in two studies, yet the studies can provide different amounts of evidence against the null hypothesis. The key is that a one-sample t test measures the difference in standard-error units, not just in the original measurement units. Sample size affects the standard error, and it also determines the degrees of freedom used to find the p-value.

This tutorial compares two hypothetical samples with the same sample mean, null value, and sample standard deviation, but with \(n=10\) and \(n=100\). Holding those summaries fixed makes it possible to see the effect of sample size clearly. The comparison builds on “The One-Sample t Test Statistic” and “Why the t Test Uses n Minus 1 Degrees of Freedom.”

Key idea: When the observed difference \(\bar{x}-\mu_0\) is not zero and, together with the sample standard deviation \(s\), stays the same, increasing \(n\) reduces the standard error \(s/\sqrt{n}\). The absolute t statistic therefore increases, and the p-value typically decreases. The degrees of freedom also increase from \(n-1\), changing the t distribution used to calculate the p-value.

What Changes When n Increases?

The one-sample t statistic divides the difference between the sample mean and the null value by the standard error. If the observed difference and \(s\) are unchanged, increasing \(n\) makes the denominator smaller. In particular, increasing the sample size from 10 to 100 multiplies \(n\) by 10 and divides the standard error by \(\sqrt{10}\). The t statistic’s absolute value is then multiplied by \(\sqrt{10}\).

$$ SE=\frac{s}{\sqrt{n}}, \qquad t=\frac{\bar{x}-\mu_0}{s/\sqrt{n}}, \qquad df=n-1. $$

The p-value depends on two things: the observed t statistic and the t distribution with the appropriate degrees of freedom. With \(n=10\), \(df=9\); with \(n=100\), \(df=99\). The distribution with 99 degrees of freedom has lighter tails than the one with 9 degrees of freedom and is closer to the standard Normal distribution. In the comparison here, both changes matter: the t statistic moves farther from 0, and its reference distribution changes.

A larger sample does not change the value of the observed mean difference in this controlled comparison. It changes how precisely that difference is estimated. A difference of 2 minutes is still 2 minutes; it represents more standard errors when the standard error is smaller.

A Careful Comparison

To isolate the sample-size effect, compare results while holding the null mean, sample mean, and sample standard deviation fixed. Then check the conditions for each sample as you would for any one-sample t test. As in “Why Conditions Matter in Mean Inference,” a calculator’s output does not replace checks of randomness, independence, and the Normal/Large Sample condition.

1
Hold the other summaries fixed.
Use the same \(\mu_0\), \(\bar{x}\), and \(s\) for the sample-size comparison.
2
Calculate each standard error and t statistic.
For each sample size, divide \(s\) by \(\sqrt{n}\), then divide \(\bar{x}-\mu_0\) by that standard error.
3
Use the matching degrees of freedom.
For each test, calculate \(df=n-1\) and use that t distribution to find the p-value for the stated alternative.
4
Compare evidence, not just means.
Compare each p-value with the same \(\alpha\), then describe the evidence about the population mean in context.

For a two-sided alternative, the p-value is the probability, assuming \(H_0\) is true, of getting a t statistic at least as far from 0 as the observed statistic. “Finding a P-Value Using tcdf” explains how the alternative determines the tail area. In the examples below, the p-values are calculated with the appropriate degrees of freedom for each sample size.

Worked Examples

Worked Example: Bottle Fill Volume With Two Sample Sizes

A fictional bottling line fills 5,000 bottles in a production run. The target mean fill volume is 500 milliliters. Consider two possible random samples from that run: one with \(n=10\), and one with \(n=100\). In both hypothetical samples, \(\bar{x}=502\) milliliters and \(s=6\) milliliters. The fill-volume distribution is approximately Normal. At \(\alpha=0.05\), compare the two-sided tests of whether the true mean differs from 500 milliliters.

State. Let \(\mu\) be the true mean fill volume, in milliliters, of bottles in this production run. For both sample-size scenarios, test \(H_0:\mu=500\) milliliters against \(H_a:\mu\ne500\) milliliters at \(\alpha=0.05\).

Plan and check conditions. Use a one-sample t test for a population mean. For each scenario, the bottles are described as randomly sampled, supporting the Random condition. For the 10% condition, the sample sizes are no more than 10% of the 5,000-bottle run: \(0.10(5000)=500\), and both \(10\leq500\) and \(100\leq500\). Thus, sampling without replacement supports independence. For \(n=10\), the large-sample route is not met, but the stated approximately Normal population supports the Normal/Large Sample condition. For \(n=100\), \(n\geq30\), so the large-sample route supports that condition. The degrees of freedom are 9 and 99, respectively.

Do. In both cases, the observed difference is \(\bar{x}-\mu_0=502-500=2\) milliliters. With \(n=10\), the standard error and t statistic are

$$ SE=\frac{6}{\sqrt{10}}\approx1.8974, \qquad t=\frac{502-500}{6/\sqrt{10}} \approx\frac{2}{1.8974} \approx1.0541, \qquad df=9. $$

For the two-sided test, the p-value is \(2P(T\geq1.0541)\) for a t distribution with 9 degrees of freedom, or about \(0.3193\). With \(n=100\), the standard error and test statistic are

$$ SE=\frac{6}{\sqrt{100}}=0.6, \qquad t=\frac{502-500}{6/\sqrt{100}} =\frac{2}{0.6} \approx3.3333, \qquad df=99. $$

For this two-sided test, the p-value is \(2P(T\geq3.3333)\) for a t distribution with 99 degrees of freedom, or about \(0.0012\). Both calculations use the same 2-milliliter difference and the same \(s=6\) milliliters. The standard error is smaller with \(n=100\), so the difference is more standard errors from the null value.

Conclude. With \(n=10\), \(0.3193>0.05\), so fail to reject \(H_0\). This sample does not provide convincing evidence that the true mean fill volume differs from 500 milliliters. With \(n=100\), \(0.0012<0.05\), so reject \(H_0\). This sample provides convincing evidence that the true mean fill volume differs from 500 milliliters. The larger sample gives stronger evidence in this comparison, even though the observed difference is the same.

Worked Example: Curing Time With a Different Mean Difference

A fictional materials lab compares the mean curing time of a product with a 40-minute reference value. Consider two possible random samples from a batch of 4,000 items. In both, the sample mean is 42 minutes and the sample standard deviation is 4 minutes. The curing-time distribution is approximately Normal. Compare two-sided tests at \(\alpha=0.05\) for \(n=10\) and \(n=100\).

State and plan. Let \(\mu\) be the true mean curing time, in minutes, for items in the batch. For each scenario, test \(H_0:\mu=40\) against \(H_a:\mu\ne40\). Use a one-sample t test. Random selection supports the Random condition. The 10% limit is \(0.10(4000)=400\), and both sample sizes are no more than 400, supporting the 10% condition. The stated approximately Normal distribution supports the Normal/Large Sample condition for \(n=10\); \(n=100\geq30\) supports the large-sample route for the other scenario.

Do. The observed difference is \(42-40=2\) minutes in both cases. For \(n=10\), \(SE=4/\sqrt{10}\approx1.2649\), so

$$ t=\frac{42-40}{4/\sqrt{10}} \approx\frac{2}{1.2649} \approx1.5811, \qquad df=9. $$

The two-sided p-value is about \(0.148\). For \(n=100\), \(SE=4/\sqrt{100}=0.4\), so

$$ t=\frac{42-40}{4/\sqrt{100}} =\frac{2}{0.4} =5.0000, \qquad df=99. $$

The two-sided p-value is less than \(0.0001\). These values come from the t distributions with 9 and 99 degrees of freedom, respectively.

Conclude. At \(\alpha=0.05\), the \(n=10\) scenario has \(p\approx0.148>0.05\), so fail to reject \(H_0\); it does not provide convincing evidence of a difference from 40 minutes. The \(n=100\) scenario has \(p<0.0001\), so reject \(H_0\); it provides convincing evidence that the true mean curing time differs from 40 minutes. Again, the mean difference is 2 minutes in both scenarios, but it is estimated more precisely in the larger sample.

Worked Example: Why Sample Size Alone Does Not Determine the Result

Suppose two fictional device-testing studies both compare mean charging time with a reference value of 60 minutes. In each study, \(\bar{x}=62\) minutes. The first has \(n=10\) and \(s=2\) minutes; the second has \(n=100\) and \(s=10\) minutes. Assume each study used a random sample from a population of at least 1,000 devices and that each population distribution is approximately Normal. Compare their two-sided tests at \(\alpha=0.05\).

State and check conditions. For each study, let \(\mu\) be the true mean charging time, in minutes, for the population of devices being studied. Test \(H_0:\mu=60\) against \(H_a:\mu\ne60\). Random sampling supports the Random condition. Since both population sizes are at least 1,000, \(n=10\) and \(n=100\) are each no more than 10% of their population sizes, supporting independence. For \(n=10\), the stated approximately Normal population supports the Normal/Large Sample condition; for \(n=100\), the large-sample route is met.

Do. In the first study, the difference is \(62-60=2\) minutes and \(SE=2/\sqrt{10}\approx0.6325\). Thus,

$$ t=\frac{62-60}{2/\sqrt{10}} \approx3.1623, \qquad df=9. $$

The two-sided p-value is about \(0.0115\). In the second study, the same 2-minute difference is paired with a larger standard deviation: \(SE=10/\sqrt{100}=1\) minute. Thus,

$$ t=\frac{62-60}{10/\sqrt{100}} =\frac{2}{1} =2.0000, \qquad df=99. $$

The two-sided p-value is about \(0.0482\).

Conclude. Both p-values are below 0.05, so both tests reject their null hypotheses. The first has the smaller p-value despite having fewer observations because its sample standard deviation is much smaller. This is not a controlled comparison isolating sample size: unlike the first examples, \(s\) changed. It shows why sample size by itself cannot predict a test result when other sample summaries differ.

Interpreting the Pattern

The first two examples hold the observed difference and sample standard deviation fixed. For \(n=10\), the standard error is \(s/\sqrt{10}\); for \(n=100\), it is \(s/10\). The latter is smaller by a factor of \(\sqrt{10}\), so the absolute t statistic is \(\sqrt{10}\) times as large. The higher degrees of freedom also change the reference distribution used to obtain the p-value.

The p-value is not the probability that \(H_0\) is true, and a smaller p-value does not mean the difference is larger in the measurement units. In the bottle example, the difference stayed at 2 milliliters. Rather, the larger sample made that same observed difference more unusual under the null model. Whether a difference matters in practice is a separate question that requires context about the measurement and its consequences.

A larger sample does not guarantee a smaller p-value in every pair of real studies. The sample means and standard deviations can differ, as they do in the device example. Even if the population mean difference is the same, a sample may show a different \(\bar{x}\) or \(s\). The controlled comparison describes what happens when the stated summaries are held fixed; it is not a promise about every new sample.

Common Mistakes and AP Exam Tips

  • Comparing only the raw mean differences. A difference of 2 units does not tell you the t statistic. Show the standard error and divide the difference by it.
  • Changing the sample size but forgetting the standard error. Recalculate \(s/\sqrt{n}\) for each \(n\); do not carry over the old denominator.
  • Using the same degrees of freedom for both tests. A one-sample t test uses \(df=n-1\), so \(n=10\) gives 9 degrees of freedom and \(n=100\) gives 99.
  • Using the wrong tail area. For \(H_a:\mu\ne\mu_0\), use a two-sided p-value. The direction of the observed difference does not turn a two-sided alternative into a one-sided test.
  • Claiming that larger samples always produce significant results. In a controlled comparison with the same nonzero difference and \(s\), a larger \(n\) increases \(|t|\). In real comparisons, however, the observed difference and \(s\) may also change.
  • Confusing statistical evidence with practical importance. State the p-value decision and evidence about the population mean in context. Do not claim that a statistically significant difference is automatically important in practice.
  • Skipping conditions because the sample is large. A large sample supports the Normal/Large Sample condition, but it does not establish random sampling or independence. Check the design and the 10% condition when sampling without replacement.
  • Writing “accept the null.” If the p-value exceeds \(\alpha\), say “fail to reject \(H_0\)” and explain that the sample does not provide convincing evidence for the alternative. This does not prove the null mean is correct.
Key takeaway: With the same nonzero \(\bar{x}-\mu_0\) and the same \(s\), increasing \(n\) from 10 to 100 reduces the standard error, increases the absolute t statistic, and typically decreases the p-value. Use \(df=n-1\) for each test, check conditions, and describe evidence about the population mean—not just the calculator output.

Check Your Understanding

Use the one-sample t test ideas from this tutorial to answer each question.

  1. A test has \(\bar{x}-\mu_0=3\) and \(s=9\). Calculate the standard error and t statistic for \(n=9\) and \(n=81\).
  2. For a one-sample t test with \(n=10\), what are the degrees of freedom? What are they when \(n=100\)?
  3. Two tests have the same nonzero mean difference and the same sample standard deviation. Explain why the test with \(n=100\) will generally have a larger absolute t statistic than the test with \(n=10\).
  4. Why must the degrees of freedom be included when comparing the two p-values?
  5. A larger sample has a smaller p-value, but the estimated mean difference is unchanged. Explain what has changed and what has not changed.
  6. Can you conclude that a larger sample always has a smaller p-value if the sample standard deviations differ? Explain briefly.