Tutorials › AP Statistics › Fail to Reject H0 Does Not Mean H0 Is True

One-sample t hypothesis tests · Tutorial 679 of 1000

Fail to Reject H0 Does Not Mean H0 Is True

Interpret large p-values cautiously: describe what the data do and do not show without treating a failure to reject as proof that the null hypothesis is true.

Intermediate 9 min read

What You'll Learn

  • Explain what a large p-value says about the evidence against a null hypothesis.
  • Distinguish “fail to reject” from accepting or proving the null hypothesis.
  • Interpret a large p-value in context for two-sided and one-sided mean tests.
  • Recognize why a test may fail to detect a difference from the null value.
  • Write conclusions that avoid claiming equality or certainty about a population mean.

A Large P-Value Is Not Proof of the Null

When a one-sample t test produces a large p-value, the data have not provided convincing evidence against the null hypothesis at the chosen significance level. That is a statement about the strength of evidence from the test. It is not a statement that the null hypothesis has been proved, or that the population mean must equal the null value.

As explained in “The Logic of a Significance Test for a Mean,” a p-value is calculated assuming \(H_0\) is true. It measures the probability of getting a test statistic at least as extreme as the observed statistic in the direction or directions specified by \(H_a\). A large p-value means that results at least this extreme are not unusual under that assumption. It does not give the probability that \(H_0\) is true.

Key distinction: If the p-value is greater than \(\alpha\), fail to reject \(H_0\). Say that the data do not provide convincing evidence for \(H_a\) at the chosen significance level. Do not say that \(H_0\) is true, has been proved, or has been accepted.

What a Large P-Value Does—and Does Not—Tell You

The decision to fail to reject means the observed data did not give a sufficiently strong signal against \(H_0\), based on the test and significance level used. It does not rule out population means other than the null value. A sample mean can be near the null value because the population mean is near it, but it can also be near the null value because of sample-to-sample variability.

The p-value is not a direct measure of how close the population mean is to \(\mu_0\). The test statistic uses the difference between the sample mean and the null value relative to the standard error, \(s/\sqrt{n}\). So the same difference in sample means can produce different test statistics in different studies. As discussed in “Effect of Sample Size on a One-Sample t Test,” sample size affects the standard error and therefore can affect how much evidence a test detects.

A large p-value also does not establish that any possible difference is unimportant. A test of \(H_0:\mu=\mu_0\) against \(H_a:\mu\ne\mu_0\) asks whether the data provide convincing evidence of a difference; failing to reject is not the same as demonstrating that the difference is small enough to be practically unimportant.

The conclusion must match the alternative hypothesis. For a two-sided alternative, describe the lack of convincing evidence that the mean differs from the null value. For a one-sided alternative, describe the lack of convincing evidence for the specific direction in \(H_a\). Do not turn that conclusion into a claim about the opposite direction unless the test was designed to examine it.

Conclusion pattern: “Because the p-value is greater than \(\alpha\), we fail to reject \(H_0\). The sample does not provide convincing evidence that [the population mean claim in \(H_a\), in context].” This wording reports the test decision without claiming the null is true.

Worked Examples

Worked Example: A Two-Sided Test Does Not Establish Equality

A fictional electronics workshop randomly selects 16 rechargeable batteries from a shipment of 400. Their mean operating time is 20.5 hours, and their sample standard deviation is 2 hours. Assume operating times in the shipment are approximately Normally distributed. At \(\alpha=0.05\), test whether the true mean operating time differs from 20 hours.

State. Let \(\mu\) be the true mean operating time, in hours, of rechargeable batteries in this shipment. Test \(H_0:\mu=20\) hours against \(H_a:\mu\ne20\) hours at \(\alpha=0.05\).

Plan and check conditions. Use a one-sample t test for a population mean. The batteries were randomly selected, supporting the Random condition. Because the sample was drawn without replacement, check the 10% condition: \(16\leq0.10(400)=40\), so independence is supported. Since \(n=16<30\), the large-sample route is not met; the stated approximately Normal population supports the Normal/Large Sample condition. The one-sample t test is appropriate.

Do. Calculate the standard error, t statistic, and degrees of freedom:

$$ SE=\frac{s}{\sqrt{n}}=\frac{2}{\sqrt{16}}=0.5\text{ hour}, \qquad t=\frac{\bar{x}-\mu_0}{s/\sqrt{n}} =\frac{20.5-20}{2/\sqrt{16}} =\frac{0.5}{0.5}=1.0000, \qquad df=16-1=15. $$

For the two-sided alternative, the p-value is the probability, assuming \(H_0\) is true, of a t statistic at least 1 unit from zero in either direction for \(df=15\). Using a t distribution calculator, the p-value is \(0.3334\), rounded. Since \(0.3334>0.05\), fail to reject \(H_0\).

Conclude. The sample does not provide convincing evidence that the true mean operating time of batteries in this shipment differs from 20 hours. This conclusion does not show that the mean is exactly 20 hours. It says the observed sample result is not unusual enough under \(H_0\) to reject the null at the 0.05 significance level.

Worked Example: A Sample Mean Equal to the Null Value Is Not Proof

A fictional school randomly selects 12 classrooms from a group of 240 to measure the average number of minutes their windows are open during a particular break. The sample mean is 7.5 minutes, and the sample standard deviation is 2.5 minutes. Assume the population distribution of these times is approximately Normal. Test at \(\alpha=0.05\) whether the true mean differs from 7.5 minutes.

State. Let \(\mu\) be the true mean number of minutes classroom windows are open during the break for classrooms in this group. Test \(H_0:\mu=7.5\) minutes against \(H_a:\mu\ne7.5\) minutes at \(\alpha=0.05\).

Plan and check conditions. Use a one-sample t test. Random selection supports the Random condition. The 10% condition is met because \(12\leq0.10(240)=24\), supporting independence. Since \(n=12<30\), the large-sample route is not met, but the stated approximately Normal population supports the Normal/Large Sample condition.

Do. The sample mean equals the null value, so the numerator of the t statistic is zero:

$$ SE=\frac{2.5}{\sqrt{12}}\approx0.7217\text{ minute}, \qquad t=\frac{7.5-7.5}{2.5/\sqrt{12}}=0.0000, \qquad df=12-1=11. $$

For a two-sided test, every possible t statistic is at least 0 units from zero. Therefore, the probability of a result at least as far from zero as the observed statistic is \(1.0000\). Since \(1.0000>0.05\), fail to reject \(H_0\).

Conclude. The sample does not provide convincing evidence that the true mean time classroom windows are open differs from 7.5 minutes. The sample mean being exactly 7.5 minutes does not prove that the population mean is 7.5 minutes. A test conclusion is about the evidence the data provide, not a declaration that a population value is known with certainty.

Worked Example: A Large P-Value for an Upper-Tailed Test

A fictional repair shop randomly selects 9 devices from a group of 180 repaired devices. The sample mean battery life after repair is 29 hours, and the sample standard deviation is 3 hours. Assume the population distribution of battery life is approximately Normal. The shop wants to know whether the true mean is greater than 30 hours. Test at \(\alpha=0.05\).

State. Let \(\mu\) be the true mean battery life, in hours, after repair for devices in this group. Test \(H_0:\mu=30\) hours against \(H_a:\mu>30\) hours at \(\alpha=0.05\).

Plan and check conditions. Use a one-sample t test. The devices were randomly selected, supporting the Random condition. The 10% condition is met because \(9\leq0.10(180)=18\), supporting independence. The sample size is \(9<30\), so the large-sample route is not met; the stated approximately Normal population supports the Normal/Large Sample condition.

Do. Calculate the standard error, test statistic, and degrees of freedom:

$$ SE=\frac{3}{\sqrt{9}}=1\text{ hour}, \qquad t=\frac{29-30}{3/\sqrt{9}}=\frac{-1}{1}=-1.0000, \qquad df=9-1=8. $$

Because \(H_a:\mu>30\), use the area to the right of \(-1.0000\) for a t distribution with 8 degrees of freedom. The p-value is \(0.8267\), rounded. Since \(0.8267>0.05\), fail to reject \(H_0\).

Conclude. The sample does not provide convincing evidence that the true mean battery life after repair is greater than 30 hours. The sample mean is below 30 hours, but this upper-tailed test does not establish that the mean is less than 30 hours. The alternative must remain the one chosen before examining the sample, and the conclusion must stay within the claim that this test examined.

Common Mistakes and AP Exam Tips

A careful conclusion separates the test decision from what the test can establish. “Fail to reject” is the decision; “the sample does not provide convincing evidence for the alternative” describes its meaning in context. Avoid using a non-significant result as if it settled the population question.

  • Writing “accept \(H_0\).” The standard decision is “fail to reject \(H_0\).” A test may not detect convincing evidence against the null without establishing the null as true.
  • Claiming the p-value is the probability that \(H_0\) is true. The p-value is calculated assuming \(H_0\) is true; it is a probability about test results under that assumption, not a probability assigned to the hypothesis.
  • Claiming the population mean equals the null value. A large p-value does not prove equality. State only that the sample does not provide convincing evidence for the alternative in context.
  • Claiming the opposite direction from a one-sided test. If \(H_a:\mu> \mu_0\), failing to reject means there is not convincing evidence for a mean greater than \(\mu_0\). It does not establish that the mean is lower.
  • Calling a large p-value evidence that a difference is unimportant. The test addresses the specified hypothesis, not whether every possible difference would matter in practice.
  • Leaving out the significance level or context. A complete conclusion reports the decision based on p and \(\alpha\), then names the population mean and the claim in the alternative hypothesis.
Key takeaway: A large p-value leads to failing to reject \(H_0\), not accepting or proving it. In context, say the data do not provide convincing evidence for the specific alternative; do not claim that the population mean equals the null value.

Check Your Understanding

For each situation, focus on what a large p-value permits you to conclude.

  1. A two-sided one-sample t test has \(p=0.27\) and \(\alpha=0.05\). State the decision and explain what the data do not establish.
  2. A student writes, “The p-value is 0.40, so there is a 40% chance that the null hypothesis is true.” Explain the error.
  3. An upper-tailed test has \(H_a:\mu>15\) and \(p=0.62\). What is an appropriate conclusion? What claim would go beyond the test?
  4. A test fails to reject \(H_0:\mu=100\). Does this demonstrate that any difference from 100 is practically unimportant? Explain.
  5. Write a context-specific conclusion for a two-sided test of whether the mean time for a fictional delivery route differs from 25 minutes, given \(p=0.31\) and \(\alpha=0.05\).