Why a Tiny Difference Can Produce Strong Evidence
In “Effect Size in Terms of the Original Units,” you learned to describe a mean difference in units that make sense in context. Here we add another piece: how much information the sample provides about that difference. When the standard deviation stays similar, a larger sample gives a more precise estimate of the mean. That precision can make a small difference produce a large t statistic and a small p-value.
The key quantity is the standard error, an estimate of how much a sample mean typically varies from sample to sample. For a one-sample t procedure, the estimated standard error of \(\bar{x}\) is \(s/\sqrt{n}\). It depends both on the sample standard deviation \(s\) and the sample size \(n\). Increasing \(n\) reduces the standard error, but not in direct proportion: multiplying \(n\) by four divides the standard error by two.
A larger magnitude of \(t\) means the observed sample mean is farther from the null value when measured in estimated standard errors. As covered in “Common Misinterpretations of the P-Value,” the p-value describes how unusual a test statistic at least this extreme would be if the null hypothesis were true. It is not a measure of how important the difference is, nor the probability that the null hypothesis is true.
This relationship explains why a huge sample can detect a difference that is small in the original units. The data may provide convincing evidence that the population mean is not exactly the null value while the estimated difference remains too small to matter for a practical decision. Statistical significance and practical importance answer different questions.
Hold the Difference Steady and Increase the Sample Size
Consider a fictional company that measures the fill volume of a product. Suppose the null mean is 500 milliliters and the sample standard deviation is about 2 milliliters. To isolate the effect of sample size, compare hypothetical random samples with the same observed difference from 500 milliliters but different sample sizes. These calculations illustrate the relationship; they do not claim that real samples would have exactly the same standard deviation and mean difference.
With a mean difference of 0.10 milliliter, a sample of 100 gives an estimated standard error of \(2/\sqrt{100}=0.20\) milliliter. The observed difference is only half a standard error from the null value. A sample of 2,500 with the same difference and standard deviation gives an estimated standard error of \(2/\sqrt{2500}=0.04\) milliliter. Now the difference is 2.5 standard errors from the null. The raw difference has not changed; its size relative to the uncertainty has.
Worked Example: The Same Small Difference in Two Samples
In two hypothetical random samples from a very large population of product fills, the sample mean is 500.10 milliliters and the sample standard deviation is 2.0 milliliters. One sample has \(n=100\); the other has \(n=2{,}500\). Test \(H_0:\mu=500\) against \(H_a:\mu\ne500\) for each sample, using \(\alpha=0.05\). Assume the measurements in each sample show no extreme outliers or severe skew.
State: Let \(\mu\) be the true mean fill volume for the population represented by these samples. The question is whether the population mean differs from 500 milliliters. In both samples, the raw estimated difference is \(500.10-500=0.10\) milliliter.
Plan: Use a one-sample t test for each sample because each compares one quantitative sample mean with a fixed benchmark. Assume each sample was randomly selected and observations are independent. For the 10% condition, the population is large enough that both 100 and 2,500 are no more than 10% of it. The samples are large, and the stated absence of extreme outliers or severe skew supports using the t procedure. Use a two-sided alternative because the question asks whether the mean differs in either direction.
Do: For \(n=100\),
The test statistic is \(t=(500.10-500)/0.20=0.50\), with \(df=100-1=99\). A two-sided t probability is about \(p=0.6182\), rounded. Since \(0.6182>0.05\), fail to reject \(H_0\).
For \(n=2{,}500\),
The test statistic is \(t=(500.10-500)/0.04=2.50\), with \(df=2{,}500-1=2{,}499\). The two-sided p-value is about \(0.0125\), rounded. Since \(0.0125<0.05\), reject \(H_0\).
Conclude: In the smaller sample, the data do not provide convincing evidence that the population mean fill volume differs from 500 milliliters. In the larger sample, the data provide convincing evidence that it differs from 500 milliliters. In both samples, however, the estimated difference is only 0.10 milliliter. Whether that change matters for product use or manufacturing depends on the context and any relevant practical benchmark; the p-value does not answer that question.
The example holds the estimated difference and sample standard deviation constant to show the mechanism clearly. In actual data, both can change from sample to sample. The standard error therefore will not always follow a perfectly predictable path in observed samples, but, with similar variability, a larger \(n\) generally makes the estimated mean more precise.
More Data Can Make a Very Small Effect Significant
A small p-value does not require a large difference in the original units. It requires a result that is unusual relative to the null model, and a small standard error can make even a small raw difference many standard errors from the null. The next example shows a tiny difference producing a very large t statistic. It also distinguishes two situations: holding the raw difference fixed as \(n\) increases, and changing the raw difference so the t statistic stays fixed.
Worked Example: A Fraction-of-a-Unit Difference
A fictional sensor company compares its average reading with a reference value of 20.0 units. Assume random sampling, independent readings, a population large enough for the 10% condition, and data without extreme outliers or severe skew. In a sample of 400 readings, suppose \(\bar{x}=20.5\) and \(s=1.0\). Test \(H_0:\mu=20.0\) against \(H_a:\mu\ne20.0\) at \(\alpha=0.05\).
State: Let \(\mu\) be the true mean sensor reading for the population represented by the random sample. The hypotheses are \(H_0:\mu=20.0\) and \(H_a:\mu\ne20.0\), measured in sensor units.
Plan: Use a one-sample t test because one sample mean is being compared with a fixed value. The sample is random, the observations are independent, and the 10% condition holds. With \(n=400\) and no extreme outliers or severe skew, the t procedure is appropriate.
Do: The estimated difference is \(20.5-20.0=0.5\) unit, and the standard error is
Thus \(t=0.5/0.05=10\), with \(df=400-1=399\). The two-sided p-value is approximately \(3.707\times10^{-21}\). This is far below \(0.05\), so reject \(H_0\).
Conclude: The data provide convincing evidence that the population mean sensor reading differs from 20.0 units. The estimated difference is 0.5 unit. Whether 0.5 unit is large enough to matter for calibration or use is a separate practical question that requires context.
Now imagine a different sample size and a smaller observed difference: \(n=1{,}600\), \(s=1.0\), and \(\bar{x}=20.25\). The standard error is \(1/\sqrt{1{,}600}=0.025\), and the difference from 20.0 is 0.25 unit. The t statistic is again \(0.25/0.025=10\), now with \(df=1{,}599\). Its two-sided p-value is approximately \(7.048\times10^{-23}\). Both results are highly statistically significant, even though the second raw difference is half as large. At the same time, these two calculations hold \(t\) fixed while changing both \(n\) and the observed difference; they are not evidence that merely increasing \(n\) makes a fixed difference produce the same t statistic.
Read the Standard Error, the Effect, and the P-Value Separately
For a fixed raw difference and similar \(s\), increasing \(n\) reduces \(SE\), increases the magnitude of \(t\), and generally reduces the p-value. But these quantities describe different things:
- The raw effect estimate gives the observed difference in the original units, such as 0.10 milliliter or 0.5 sensor unit.
- The standard error describes the estimated sample-to-sample variability of the sample mean under the sampling model. It is not the standard deviation of individual observations.
- The t statistic measures the estimated difference from the null value in standard-error units.
- The p-value describes how unusual a result at least as extreme as the observed test statistic would be if the null hypothesis were true.
For two independent samples, the same general idea applies. The estimated standard error of the difference in sample means is based on both groups’ variability and sample sizes: \(\sqrt{s_1^2/n_1+s_2^2/n_2}\). Increasing the group sample sizes usually reduces this standard error when the group variability stays similar. The raw difference \(\bar{x}_1-\bar{x}_2\) still needs to be interpreted in the response’s original units and with the group order stated, as discussed in the earlier tutorials on choosing mean procedures and describing effect size.
A practical benchmark helps keep the interpretation grounded. If a difference of 0.10 milliliter would not affect product quality or a decision, a small p-value alone does not make it consequential. Conversely, a difference that matters in a particular setting is not rendered unimportant by a larger p-value. Consider the effect’s size and context alongside the evidence and uncertainty.
Common Mistakes and AP Exam Tips
- Calling the standard error the standard deviation: The standard deviation \(s\) describes spread among individual observations. The standard error \(s/\sqrt{n}\) describes estimated variability of the sample mean. State which quantity you mean.
- Claiming that larger samples change the raw effect: Larger \(n\) does not automatically increase \(\bar{x}-\mu_0\). It can make the same estimated difference more precise, so the test statistic changes even if the raw difference does not.
- Equating a small p-value with practical importance: A small p-value is evidence against the null hypothesis under the test model. It does not say that the effect is large or useful. Report the difference in its original units and discuss a relevant context benchmark.
- Using “accept the null” after a large p-value: As emphasized in “Fail to Reject H0 Wording That Earns Full Credit,” say “fail to reject \(H_0\)” and explain that the data do not provide convincing evidence for the alternative. Do not claim the null has been proved.
- Omitting conditions because the sample is huge: A large sample does not fix a biased sampling method, dependent observations, or a violated 10% condition. Identify the random process, independence, 10% condition when relevant, and data-shape requirements for the chosen t procedure.
A complete AP response connects the p-value decision to the population parameter and the question’s context. For example: “Because \(p\approx0.0125<0.05\), reject \(H_0\). The data provide convincing evidence that the population mean fill volume differs from 500 milliliters. The estimated difference is only 0.10 milliliter, so its practical importance depends on whether a change of that size matters for the product.” This distinguishes evidence from effect size rather than letting one stand in for the other.
Check Your Understanding
Use the relationship between sample size, standard error, and the t statistic to answer each question.
- If \(s=6\), what is the estimated standard error for a sample mean when \(n=36\)? What would the standard error be if \(n\) increased to 144 and \(s\) stayed the same?
- A sample mean differs from the null value by 0.3 unit. If \(s\) stays about the same and \(n\) increases, what happens to the estimated standard error and the magnitude of the t statistic?
- Why does a small p-value from a very large sample not, by itself, show that a difference matters in practice?
- For a one-sample t test with \(n=80\), \(\bar{x}=5.2\), and \(\mu_0=5.0\), what additional sample information is needed to calculate the test statistic?
- A test has \(p=0.18\) at \(\alpha=0.05\). Give the correct decision wording and explain why it does not prove that the population mean equals the null value.