Tutorials › AP Statistics › Why a t Test Never Proves the Null Mean

P-values and mean-inference conclusions · Tutorial 750 of 1000

Why a t Test Never Proves the Null Mean

Use t tests and confidence intervals for means to distinguish failing to find a difference from showing that plausible differences are small.

Intermediate 9 min read

What You'll Learn

  • Explain why failing to reject a null hypothesis does not prove the null mean is true.
  • Connect a nonsignificant one-sample t test to a confidence interval that includes the null value.
  • Use a narrow confidence interval to discuss which mean values remain plausible.
  • Distinguish exact equality from evidence against departures large enough to matter in context.
  • Write careful conclusions for one-sample and two-sample t procedures.

Not Finding a Difference Is Not Proving Equality

In “Fail to Reject H0 Wording That Earns Full Credit,” you learned that when the p-value is greater than \(\alpha\), you fail to reject \(H_0\). That decision means the data do not provide convincing evidence for the alternative claim at the chosen significance level. It does not prove that the null hypothesis is true.

This distinction matters in t procedures for population means. A sample mean may be close to the value in \(H_0\), or a test may have a p-value above \(\alpha\), without establishing that the population mean equals that value exactly. Sample results vary, and a test can fail to detect a difference even when the true mean is not the null value.

Key distinction: “The data do not provide convincing evidence that the mean differs from the null value” describes a failure to reject a two-sided null hypothesis. “The mean equals the null value” is a claim of exact equality that the test has not proved.

A confidence interval adds information about which population mean values are reasonably compatible with the sample, under the assumptions of the procedure. If the interval includes the null value, that is consistent with failing to reject the matching two-sided t test. But an interval containing the null value does not prove the null value is correct. The interval also contains other values that remain plausible.

The width of the interval is important. A wide interval may include the null value and many substantially different means, leaving considerable uncertainty. A narrow interval that includes the null value may rule out large departures from it while still not establishing exact equality. That is the difference between absence of evidence for a difference and more informative evidence that differences of a specified size are unlikely.

What the Test and Interval Can Tell You

For a one-sample t test of \(H_0:\mu=\mu_0\) against \(H_a:\mu\ne\mu_0\), the test assesses whether the sample provides convincing evidence that the true mean differs from \(\mu_0\). When the test fails to reject, it has not shown that \(\mu=\mu_0\); the result may reflect a small difference, substantial sample variability, or both.

A confidence interval helps show the range of means compatible with the data. As covered in “Linking Two-Sample Intervals and Tests,” for the corresponding two-sided test and interval, the null value is included in the interval when the test does not reject at the matching significance level. This connection applies to one-sample t procedures as well: a 95% one-sample t interval corresponds to a two-sided test at \(\alpha=0.05\).

How to read a nonsignificant result: First say that the data do not provide convincing evidence for the alternative claim. Then inspect the confidence interval: does it include many meaningfully different values, or is it narrow enough to rule out departures that would matter in context? Neither result proves exact equality.

To discuss whether a difference would matter, the context must provide a meaningful benchmark for the size of a difference. For example, a process might treat a change of one milliliter as important. If a confidence interval lies entirely within one milliliter above or below the target, it gives more precise information than simply reporting a p-value above 0.05. This comparison is not a proof that the mean equals the target, and it is not, by itself, a formal equivalence test. It is a contextual reading of the values in the interval.

Worked Example: Mean Time at a Community Workshop

Worked Example: Mean Time at a Community Workshop

A fictional community workshop expects each repair appointment to last 60 minutes on average. Organizers take a random sample of 16 appointments. The sample mean is 62 minutes, and the sample standard deviation is 8 minutes. The sample is less than 10% of the relevant appointments, and the appointment times show no severe skewness or extreme outliers. Use \(\alpha=0.05\) to test whether the true mean appointment time differs from 60 minutes.

1
State.
Let \(\mu\) be the true mean appointment time, in minutes, for repair appointments at this workshop during the period represented by the sample. The hypotheses are \(H_0:\mu=60\) minutes and \(H_a:\mu\ne60\) minutes.
2
Plan and check conditions.
Use a one-sample t test for a population mean. The appointments were randomly sampled, meeting the random condition. The sample is less than 10% of the relevant appointments, supporting independence under the 10% condition. There are no severe skewness or extreme outliers, so the sample data are reasonably compatible with a t procedure. The two-sided alternative matches the question about a difference in either direction.
3
Do.
The degrees of freedom are \(16-1=15\). The standard error is \(s/\sqrt{n}=8/\sqrt{16}=8/4=2\) minutes. The test statistic is
\(t=(\bar{x}-\mu_0)/(s/\sqrt{n})=(62-60)/(8/\sqrt{16})=2/2=1.00.\)
For a two-sided test with 15 degrees of freedom, the p-value is approximately \(0.3333\), rounded to four decimal places. Since \(0.3333>0.05\), fail to reject \(H_0\). A 95% t interval uses \(t^*\approx2.131\):
\(62\pm2.131(2)=62\pm4.262,\)
or approximately \((57.74,66.26)\) minutes.
4
Conclude in context.
The data do not provide convincing evidence that the true mean repair appointment time at this workshop differs from 60 minutes. The confidence interval includes 60 minutes, but it also includes other values, from about 57.74 to 66.26 minutes. The test and interval therefore do not establish that the true mean is exactly 60 minutes.

The interval includes values up to about 2.26 minutes below the target and 6.26 minutes above it, so departures of more than four minutes above the target are still included, but departures of more than four minutes below it are not. So the result is not just a failure to detect a difference; it also shows that this sample does not estimate the mean precisely enough to exclude several potentially meaningful departures from 60 minutes.

A Narrow Interval Gives More Information

A failure to reject \(H_0\) can accompany very different levels of precision. In the workshop example, the interval was wide. In the next example, the interval is much narrower. Comparing the interval with a context-based benchmark helps explain what the data can rule out and what they cannot.

Worked Example: Filling Reusable Water Bottles

A fictional production line aims to fill reusable bottles to a mean of 500 milliliters. A random sample of 100 bottles has a mean fill of 500 milliliters and a standard deviation of 4 milliliters. The sample is less than 10% of the relevant bottles, and a plot shows no severe skewness or extreme outliers. Test \(H_0:\mu=500\) against \(H_a:\mu\ne500\) at \(\alpha=0.05\), and calculate a 95% confidence interval.

Plan and conditions: Use a one-sample t test and t interval. The bottles were randomly sampled, so the random condition is met. The sample is less than 10% of the relevant bottles, supporting independence under the 10% condition. With \(n=100\) and no severe skewness or extreme outliers, a t procedure is reasonable.

Do: The standard error is \(s/\sqrt{n}=4/\sqrt{100}=4/10=0.4\) milliliters. The test statistic is

$$ t=\frac{\bar{x}-\mu_0}{s/\sqrt{n}} =\frac{500-500}{4/\sqrt{100}} =0. $$

With \(df=100-1=99\), the two-sided p-value is \(1.0000\). Since \(1.0000>0.05\), fail to reject \(H_0\). For a 95% interval, \(t^*\approx1.984\), so

$$ 500\pm1.984(0.4) =500\pm0.7936 \approx(499.21,500.79)\text{ milliliters}. $$

Conclude: The data do not provide convincing evidence that the true mean fill differs from 500 milliliters. The interval is narrow and includes values only about 0.79 milliliters below or above the target. If a departure of one milliliter or more would matter operationally, the interval gives useful evidence against departures that large at the 95% confidence level. It still does not prove that the true mean is exactly 500 milliliters.

Here, the sample mean happens to equal the target, and the test statistic is zero. That explains why the p-value is 1.0000: the observed statistic is not extreme under the null. But even this especially close sample result is not proof of the null mean. The interval describes uncertainty about the population mean, not certainty that the mean equals its center.

Worked Example: Comparing Two Neighborhoods

Worked Example: Comparing Two Neighborhoods

A fictional planning group compares weekly hours of outdoor activity for residents in two neighborhoods. Independent random samples give \(n_1=20\), \(\bar{x}_1=12\) hours, and \(s_1=5\) hours for neighborhood 1; and \(n_2=20\), \(\bar{x}_2=10\) hours, and \(s_2=5\) hours for neighborhood 2. Each sample is less than 10% of its neighborhood population. Neither group’s data show severe skewness or extreme outliers. At \(\alpha=0.05\), test whether the true mean weekly hours differ, and form a 95% confidence interval for the difference.

State: Let \(\mu_1\) and \(\mu_2\) be the true mean weekly outdoor-activity times, in hours, for residents of neighborhoods 1 and 2, respectively. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\).

Plan and check conditions: Use an unpooled two-sample t test and interval. The two samples are independent random samples, meeting the random-design and independent-groups requirements. Each sample is less than 10% of its respective neighborhood population, supporting the 10% condition for both groups. Neither group shows severe skewness or extreme outliers, so a t procedure is reasonable for each group. The two-sided alternative matches the question about a difference in either direction.

Do: The observed difference is \(\bar{x}_1-\bar{x}_2=12-10=2\) hours. The standard error is

$$ \sqrt{\frac{s_1^2}{n_1}+\frac{s_2^2}{n_2}} =\sqrt{\frac{25}{20}+\frac{25}{20}} =\sqrt{2.5} \approx1.5811\text{ hours}. $$

Thus \(t=2/1.5811\approx1.2649\). The Welch degrees of freedom are 38, and the two-sided p-value is approximately \(0.2136\), rounded to four decimal places. Since \(0.2136>0.05\), fail to reject \(H_0\). For 38 degrees of freedom, \(t^*\approx2.024\). The 95% interval is

$$ 2\pm2.024(1.5811) \approx2\pm3.20 \approx(-1.20,5.20)\text{ hours}. $$

Conclude: The data do not provide convincing evidence that the true mean weekly outdoor-activity times differ between residents of these neighborhoods. The interval includes zero, but it also includes differences as large as about 5.20 hours in the direction of neighborhood 1 and about 1.20 hours in the opposite direction. Thus the study has not established equal means; the interval also shows substantial uncertainty about the size and direction of a possible difference.

Common Mistakes and AP Exam Tips

  • Writing “accept \(H_0\)” or “prove \(H_0\)”: A t test either rejects \(H_0\) or fails to reject it. A full-credit conclusion says the data do not provide convincing evidence for the claim in \(H_a\), in context.
  • Turning a nonsignificant result into equality: A p-value above \(\alpha\) does not show that the means are exactly equal. Say that the test did not provide convincing evidence of a difference.
  • Ignoring the confidence interval’s width: An interval containing the null value may be wide or narrow. Report what other values it includes before describing how informative it is.
  • Calling a p-value the probability that the null is true: A p-value is calculated assuming \(H_0\) is true; it is not the probability that \(H_0\) is true.
  • Claiming a precise interval proves exact equality: A narrow interval can make large departures less plausible, especially relative to a meaningful benchmark. It still does not identify the exact population mean.
  • Forgetting the context and units: State which population mean or difference in means is being discussed, and include units when they clarify the result.

For a two-sided test, a careful conclusion might read: “Since the p-value is greater than \(\alpha=0.05\), fail to reject \(H_0\). The data do not provide convincing evidence that the true mean [or the two true population means] differs [differ] from the null value in this context.” Then use the interval to describe the range and precision of plausible mean values.

Key takeaway: A t test that fails to reject the null mean has not proved it. A confidence interval shows which mean values remain plausible: a wide interval signals substantial uncertainty, while a narrow interval may rule out practically important departures without proving exact equality.

Check Your Understanding

For each situation, distinguish what the test decision says from what the interval says about plausible population means.

  1. A one-sample t test of \(H_0:\mu=25\) against \(H_a:\mu\ne25\) has \(p=0.18\). What is the appropriate decision and contextual conclusion if \(\mu\) is the mean battery life, in hours, for a specified device model?
  2. A 95% t interval for a population mean is \((48,61)\), and the null value is 50. Why does including 50 not prove that the population mean is 50?
  3. Two bottling-line means are compared, and the test fails to reject equality. What additional information does a confidence interval provide that the p-value alone does not?
  4. A 95% interval is narrow and lies within a prespecified range of differences considered operationally unimportant. What can this suggest about larger departures, and what can it not prove?
  5. Explain why “the sample mean equals the null value” is not the same claim as “the population mean equals the null value.”