When a Real Difference Is Hard to Detect
In “Why Huge Samples Make Tiny Differences Significant,” you saw how a large sample can make a small difference stand out from the sample-to-sample variability. The other side of that idea is just as important: a small sample can leave substantial uncertainty. Even when a population difference really exists, a sample may not show it clearly, and a test may produce a large p-value.
A t statistic measures an observed difference in estimated standard-error units. When the sample is small, the standard error can be large, so a real difference may be only a modest number of standard errors from the null value. A modest t statistic can correspond to a large p-value. That result does not show that the null hypothesis is true; it says the data do not provide strong evidence against it.
A large p-value is not the probability that the null hypothesis is true, nor is it the probability that the result is a Type II error. It is a probability calculated under the assumption that the null hypothesis is true. As covered in “Fail to Reject H0 Wording That Earns Full Credit,” the appropriate conclusion is that the data do not provide convincing evidence for the alternative hypothesis—not that the null has been proved.
Why the Standard Error Matters
For a one-sample t test, the standard error of the sample mean is \(s/\sqrt{n}\). For a paired t test, the same formula applies to the sample standard deviation of the pairwise differences. For a two-sample t test, the estimated standard error of the difference between sample means depends on the variability and sample size in both groups. In each case, a larger standard error means that the observed difference is less striking relative to the uncertainty in the estimate.
This is not a rule that every small-sample test has a large p-value. A small sample can produce a small p-value if the observed difference is large compared with the variability. Nor does a large p-value tell us that the effect is unimportant. Statistical evidence, the estimated effect in its original units, and practical importance are related but distinct considerations.
Worked Examples: A Real Difference Can Be Missed
Worked Example: Battery Runtime in a Small Sample
Suppose, for this teaching example, that the actual mean runtime of a certain type of battery is 52 hours. A researcher does not know this value and tests whether the population mean differs from the 50-hour benchmark. A random sample of 15 batteries has \(\bar{x}=52\) hours and \(s=1.5\sqrt{15}\approx5.809\) hours. Assume the sample’s runtime distribution has no extreme outliers or severe skew, and test at \(\alpha=0.05\).
State: Let \(\mu\) be the true mean runtime, in hours, for the population represented by the sample. The hypotheses are \(H_0:\mu=50\) and \(H_a:\mu\ne50\). In this hypothetical example, the actual value is 52 hours, so the null hypothesis is false.
Plan: Use a one-sample t test because one quantitative sample mean is being compared with a fixed benchmark. The batteries are assumed to have been randomly sampled, and the observations are independent. The population is large enough that 15 batteries are no more than 10% of it, satisfying the 10% condition. With \(n=15\), the stated absence of extreme outliers or severe skew supports using the t procedure. A two-sided alternative is appropriate because the question asks whether the mean differs in either direction.
Do: The estimated standard error is
The test statistic is
The degrees of freedom are \(15-1=14\). For a two-sided t test with \(t=1.3333\) and 14 degrees of freedom, \(p\approx0.2037\), rounded. Since \(0.2037>0.05\), fail to reject \(H_0\).
Conclude: The sample does not provide convincing evidence that the population mean battery runtime differs from 50 hours. In this teaching example, we stipulated that the actual mean is 52 hours, so this outcome is a Type II error. The researcher could not determine that from this sample’s large p-value alone.
The sample mean in this example happens to be two hours above the benchmark, but the estimated standard error is 1.5 hours. That makes the observed difference only about 1.33 standard errors from the null value. The large p-value reflects how compatible a result at least this extreme is with the null model—not whether the stipulated actual mean is known to the researcher.
Worked Example: Comparing Two Small Groups
A fictional school tests two study programs using independent random samples of students. The first group has \(n_1=8\), a mean score of 82 points, and a standard deviation of 6 points. The second group has \(n_2=8\), a mean score of 78 points, and a standard deviation of 6 points. For illustration, suppose the actual population means differ by 4 points. Test whether the population means differ at \(\alpha=0.05\).
State: Let \(\mu_1\) and \(\mu_2\) be the population mean scores for students represented by the first and second samples, respectively. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\). The sample estimate, defined as first group minus second group, is \(82-78=4\) points.
Plan: Use an unpooled two-sample t test because the response is quantitative and the two samples are independent rather than paired. Assume the samples were randomly selected, each population has at least 80 students so each sample is no more than 10% of its population, and observations within and between the samples are independent. The score distributions in both groups show no extreme outliers or severe skew. These conditions support the two-sample t procedure. Use a two-sided alternative because the question asks whether the means differ.
Do: The estimated standard error for the difference in sample means is
Thus,
The Welch degrees of freedom are 14 because the two sample sizes and sample standard deviations are equal. The two-sided p-value is approximately \(0.2037\), rounded. Since \(0.2037>0.05\), fail to reject \(H_0\).
Conclude: The data do not provide convincing evidence that the population mean scores differ between students represented by the two groups. In this hypothetical setup, the population means really do differ by 4 points; nevertheless, the small samples and the observed variation leave the estimated difference insufficiently far from zero, in standard-error units, to produce a small p-value. The test has not proved that the population means are equal.
The second example shows why a raw difference should not be read in isolation. Four points may be noticeable on a particular score scale, but the test weighs that estimate against the uncertainty in the comparison. The p-value summarizes evidence against the null hypothesis; it does not give the chance that the 4-point population difference is real.
Worked Example: A Small Paired Sample of Plant Growth
A gardener measures the change in height, in centimeters, for each of four randomly selected plants after trying a new watering schedule. Define each difference as height after the schedule minus height before it. The sample mean difference is \(\bar{x}_d=2\) centimeters and the standard deviation of the differences is \(s_d=4\) centimeters. For illustration, suppose the actual population mean change is 2 centimeters. Test whether the schedule changes mean plant height at \(\alpha=0.05\).
State: Let \(\mu_d\) be the true mean change in height for the population represented by the plants, using the stated after-minus-before order. Test \(H_0:\mu_d=0\) against \(H_a:\mu_d\ne0\).
Plan: Use a paired t test because each plant is measured twice. The analysis uses one difference per plant, not two independent groups of heights. Assume the plants are a random sample, the differences from distinct plants are independent, and the population is large enough that four plants are no more than 10% of it. The plot of the four differences is described as roughly symmetric with no outliers. This shape assessment is especially important because the sample is small. Use a two-sided alternative because the question asks whether the schedule changes mean height in either direction.
Do: The standard error of the sample mean difference is
The test statistic and degrees of freedom are
The two-sided p-value for \(t=1\) with 3 degrees of freedom is approximately \(0.3910\), rounded. Since \(0.3910>0.05\), fail to reject \(H_0\).
Conclude: The data do not provide convincing evidence that the population mean change in plant height differs from zero. In this illustration the actual mean change is stipulated to be 2 centimeters, so failing to reject is a Type II error. The large p-value is not proof that the watering schedule has no effect.
What a Large P-Value Does—and Does Not—Tell You
Across the examples, the differences were not necessarily zero, but they were modest relative to their estimated standard errors. The resulting test statistics were not far enough into the tails of their t distributions to produce small p-values. With small samples, estimates often have substantial sampling variability, so a real effect can be hard to distinguish from the variation expected under the null model.
In ordinary research, a test cannot reveal whether a particular large-p-value result is a Type II error, because the population parameter is unknown. What you can report is the evidence decision and its limits. If \(p>\alpha\), fail to reject the null hypothesis and say that the data do not provide convincing evidence for the alternative in context. Do not say that the null has been accepted, confirmed, or proved.
A large p-value also does not guarantee that the true effect is large, small, or zero. It indicates that the observed data are not unusual enough under the null model to provide convincing evidence against that model at the chosen significance level. The size of the estimated difference in original units remains useful context, as emphasized in “Effect Size in Terms of the Original Units.”
Common Mistakes and AP Exam Tips
- Claiming that a large p-value proves no difference: It does not. Full-credit wording says “fail to reject \(H_0\)” and explains that the data do not provide convincing evidence for \(H_a\), in context.
- Calling the p-value the chance the null is true: A p-value is calculated assuming the null hypothesis is true. It is not the probability that the null is true or that the observed conclusion is wrong.
- Assuming every small-sample study must have a large p-value: Sample size is only one contributor to uncertainty. A large observed effect relative to variability can still produce strong evidence in a small sample.
- Confusing an actual Type II error with what the researcher knows: If the null is truly false and the test fails to reject it, that outcome is a Type II error. In a real study, the truth is typically unknown, so a large p-value alone cannot establish that such an error occurred.
- Skipping conditions because the example is simple: Name the random sampling or assignment process, independence, the 10% condition when sampling without replacement, and the relevant data-shape condition for the chosen t procedure.
A strong AP response separates the calculation from the conclusion. Report the test statistic and p-value, compare the p-value with the stated \(\alpha\), and then explain what the evidence does and does not support about the population parameter. If discussing a stipulated true difference in a teaching example, identify that assumption clearly rather than implying a real study could know the population truth from its sample.
Check Your Understanding
Use the relationship between sample size, variability, and the t statistic to answer each question.
- A one-sample t test gives \(p=0.27\) at \(\alpha=0.05\). State the decision and its meaning in context, without claiming that the null hypothesis has been proved.
- Why can a sample mean that differs from the null value still produce a large p-value when the sample is small?
- In a hypothetical setting, the population mean is known to differ from the null value, but a test fails to reject. What is this outcome called?
- For a paired t test, what observations should be checked for strong skew or outliers: the original measurements or the pairwise differences?
- Does a large p-value tell you the probability that the null hypothesis is true? Explain briefly.