Tutorials › AP Statistics › Fail to Reject H0 Wording That Earns Full Credit

P-values and mean-inference conclusions · Tutorial 749 of 1000

Fail to Reject H0 Wording That Earns Full Credit

Practice turning a non-significant t-test result into a contextual conclusion without treating it as proof that the null hypothesis is true.

Intermediate 9 min read

What You'll Learn

  • Compare a p-value with alpha and state the correct fail-to-reject decision.
  • Explain what a non-significant p-value says under the null hypothesis.
  • Write a contextual conclusion that matches the alternative hypothesis.
  • Distinguish insufficient evidence from evidence that population means are equal.
  • Identify common wording that overstates what a t test can establish.

What “Fail to Reject” Means

In “Reject H0 Wording That Earns Full Credit,” you learned to connect a rejection decision to evidence for the alternative hypothesis. When the p-value is greater than the significance level \(\alpha\), the decision is different: fail to reject \(H_0\). The wording matters because this decision does not establish that the null hypothesis is true.

A p-value greater than \(\alpha\) means the observed result is not sufficiently unusual under \(H_0\), by the decision rule chosen for the test, to count as convincing evidence against \(H_0\). The data may still be compatible with values described by \(H_a\); the test simply has not provided convincing evidence for that alternative at the selected significance level.

Conclusion pattern: Since the p-value is greater than \(\alpha\), fail to reject \(H_0\). The data do not provide convincing evidence that [state the claim in \(H_a\), in context]. This conclusion is not evidence that \(H_0\) is true or that the population means are exactly equal.

The final sentence should describe the population parameter in the alternative hypothesis. For a one-sided test, state that the data do not provide convincing evidence that the true mean is greater than or less than the target, as appropriate. For a two-sided test, state that the data do not provide convincing evidence that the true mean differs from the target. For two means, name the two populations and the measured variable.

This is a conclusion about the strength of evidence from the test, not a declaration that nothing is happening. A study may fail to detect a real difference, for example, if the difference is modest or the sample results are variable. The test result alone does not tell you that the null value is correct.

Build a Clear Fail-to-Reject Conclusion

As in “Making a Decision From P-Value and Alpha,” compare the p-value with the significance level selected for the test. If \(p>\alpha\), fail to reject \(H_0\). Then use the alternative hypothesis to state what the data do not provide convincing evidence for.

1
State the decision.
Write “Since the p-value is greater than \(\alpha\), fail to reject \(H_0\).” Do not switch to “accept \(H_0\).”
2
Explain the p-value, if useful.
Describe it as the probability, assuming \(H_0\) is true, of getting a test statistic at least as extreme as the observed statistic in the direction or directions specified by \(H_a\).
3
Give the contextual conclusion.
Say that the data do not provide convincing evidence for the claim in \(H_a\). Name the population mean or means and the variable, including units when useful.

The p-value explanation describes how the result compares with what would be expected under the null hypothesis. The contextual conclusion answers the research question with appropriately cautious language. A short answer may not need to repeat the full definition of the p-value, but it should make the decision and the implication for the alternative clear.

A useful way to check your final sentence is to compare it directly with \(H_a\). If \(H_a\) claims a mean is below a target, the conclusion should say that the data do not provide convincing evidence that the population mean is below the target. Do not replace that with a claim that the mean is above the target; failure to find evidence in one direction does not establish the opposite direction.

Worked Example: A Clinic’s Appointment Wait Time

Worked Example: A Clinic’s Appointment Wait Time

A fictional clinic wants to know whether its mean weekday appointment wait is less than 20 minutes. Staff select a random sample of 25 weekday appointments. The sample mean wait is 18.5 minutes, and the sample standard deviation is 6 minutes. Assume the appointments are less than 10% of the clinic’s relevant weekday appointments, and the sample data show no severe skewness or extreme outliers. Use \(\alpha=0.10\).

1
State.
Let \(\mu\) be the true mean weekday appointment wait, in minutes, for appointments at this clinic during the period represented by the study. The hypotheses are \(H_0:\mu=20\) minutes and \(H_a:\mu<20\) minutes.
2
Plan and check conditions.
Use a one-sample t test for a population mean. The clinic used a random sample, so the random condition is met. The sample is less than 10% of the relevant appointments, supporting independence under the 10% condition. With no severe skewness or extreme outliers, using a t procedure is reasonable. The left-sided alternative matches the question about waits below 20 minutes.
3
Do.
The degrees of freedom are \(25-1=24\). The standard error is
\(\dfrac{s}{\sqrt{n}}=\dfrac{6}{\sqrt{25}}=\dfrac{6}{5}=1.2\) minutes.
The test statistic is
\(t=\dfrac{\bar{x}-20}{s/\sqrt{n}}=\dfrac{18.5-20}{6/\sqrt{25}}=\dfrac{-1.5}{1.2}=-1.25.\)
For a left-tailed test with 24 degrees of freedom, the p-value is \(P(T\le-1.25)\approx0.1117\), rounded to four decimal places. Since \(0.1117>0.10\), fail to reject \(H_0\).
4
Conclude in context.
If the true mean weekday wait were 20 minutes, the probability of getting a t statistic of \(-1.25\) or lower would be about \(0.1117\). Since the p-value is greater than \(\alpha=0.10\), the data do not provide convincing evidence that the true mean weekday appointment wait at this clinic is less than 20 minutes.

The conclusion does not say that the true mean wait is 20 minutes. It says that this sample and test do not provide convincing evidence for the specific claim that the mean is below 20 minutes at the chosen significance level. The sample mean of 18.5 minutes is below 20, but that sample result alone does not establish the population claim.

Match the Wording to the Alternative

A fail-to-reject conclusion should negate the claim in \(H_a\) as a claim that the data have not established—not assert the opposite claim. Use the parameter and direction in the hypotheses to keep the meaning precise.

Alternative hypothesisCareful fail-to-reject conclusion
\(H_a:\mu<\mu_0\)The data do not provide convincing evidence that the true population mean is less than \(\mu_0\).
\(H_a:\mu>\mu_0\)The data do not provide convincing evidence that the true population mean is greater than \(\mu_0\).
\(H_a:\mu\ne\mu_0\)The data do not provide convincing evidence that the true population mean differs from \(\mu_0\).
\(H_a:\mu_1-\mu_2<0\)The data do not provide convincing evidence that the true mean for population 1 is less than the true mean for population 2.
\(H_a:\mu_1-\mu_2\ne0\)The data do not provide convincing evidence that the two true population means differ.

For example, if the alternative is two-sided, a fail-to-reject decision does not prove the means are equal. It means the test did not provide convincing evidence of a difference at the selected \(\alpha\). Likewise, for a one-sided alternative, do not conclude the population mean must lie on the other side of the null value.

Worked Example: Testing Whether Seedlings Take Longer to Mature

A fictional greenhouse compares a new growing mix with a 30-day maturity benchmark. The question is whether the true mean time to maturity for seedlings using the new mix is greater than 30 days. A one-sample t test is conducted on a random sample of 32 seedlings, and the reported p-value is \(0.084\). The sample is less than 10% of the relevant seedlings; the data show no severe skewness or extreme outliers. Use \(\alpha=0.05\).

State: Let \(\mu\) be the true mean time to maturity, in days, for seedlings grown with the new mix under the conditions represented by the study. The hypotheses are \(H_0:\mu=30\) days and \(H_a:\mu>30\) days.

Plan: Use a one-sample t test. The seedlings were randomly sampled, meeting the random condition. The sample is less than 10% of the relevant population, supporting independence under the 10% condition. The data do not show severe skewness or extreme outliers, so a t procedure is reasonable. A right-tailed test matches the claim that the mean maturity time is greater than 30 days.

Do: The reported p-value, \(0.084\), is greater than \(\alpha=0.05\), so fail to reject \(H_0\). Assuming the true mean maturity time is 30 days, the probability of obtaining a test statistic at least as large as the observed one in the right-tail direction is \(0.084\).

Conclude: The data do not provide convincing evidence that the true mean time to maturity for seedlings grown with the new mix is greater than 30 days. This is not evidence that the true mean is exactly 30 days or that it is less than 30 days.

Worked Example: Comparing Weekly Biking Times

A fictional community group compares weekly biking time for residents of two neighborhoods. Independent random samples of 25 and 27 residents are selected. The study uses an unpooled two-sample t test to assess whether the population mean weekly biking times differ. The reported p-value is \(0.28\). Each sample is less than 10% of its neighborhood population, and neither group’s data show severe skewness or extreme outliers. Use \(\alpha=0.05\).

State: Let \(\mu_1\) be the true mean weekly biking time, in hours, for residents of neighborhood 1, and let \(\mu_2\) be the corresponding true mean for residents of neighborhood 2. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\).

Plan: Use an unpooled two-sample t test. The samples are independent random samples, satisfying the random-design requirement and the independent-groups condition. Each sample is less than 10% of its respective neighborhood population, supporting the 10% condition for both groups. Neither group shows severe skewness or extreme outliers, so the data are reasonably compatible with a t procedure. The two-sided alternative is appropriate because the question asks whether the means differ in either direction.

Do: Since \(0.28>0.05\), fail to reject \(H_0\). Assuming the two population mean biking times are equal, the probability of getting a test statistic at least as extreme as the observed one in either direction is \(0.28\).

Conclude: The data do not provide convincing evidence that the true mean weekly biking times differ between residents of these two neighborhoods. The conclusion is about evidence for a difference; it does not establish that the two population means are equal.

Common Mistakes and AP Exam Tips

  • Writing “accept \(H_0\)”: A test that fails to reject the null does not confirm it. Write “fail to reject \(H_0\)” and explain that the data do not provide convincing evidence for the alternative claim.
  • Claiming the means are equal: Failing to find convincing evidence of a difference is not the same as proving equality. State that the test did not provide convincing evidence that the means differ.
  • Claiming the opposite of \(H_a\): If \(H_a\) says the mean is greater than a target and the test fails to reject, that does not establish that the mean is less than or equal to the target. Stick to what the test did not provide evidence for.
  • Describing only the sample: “The sample mean was below the target” reports a statistic, not the inferential conclusion. Name the true population mean and the context.
  • Calling the p-value the probability that \(H_0\) is true: The p-value is calculated assuming \(H_0\) is true. It is a probability about test statistics under that assumption, not a probability that a hypothesis is true.
  • Using “no effect” or “no difference” without qualification: Those phrases can sound like claims of certainty. Prefer “the data do not provide convincing evidence of a difference” when the alternative is two-sided.
  • Ignoring the study design: State conclusions about the population the study represents. Do not claim a treatment caused an outcome unless the design supports a causal conclusion, such as through random assignment in an experiment.

A concise, full-credit response can be: “Since the p-value is ___ and \(\alpha\) is ___, fail to reject \(H_0\). The data do not provide convincing evidence that [the contextual claim in \(H_a\)].” Add a p-value explanation when it helps show your reasoning, but do not turn a non-significant result into acceptance of the null.

Key takeaway: When \(p>\alpha\), fail to reject \(H_0\). Conclude that the data do not provide convincing evidence for the claim in \(H_a\), in context. Do not claim that \(H_0\) is true, that population means are exactly equal, or that the opposite of \(H_a\) has been established.

Check Your Understanding

For each situation, focus on the decision and on what the data do—and do not—support in context.

  1. A one-sample t test has \(p=0.12\) and \(\alpha=0.05\), with \(H_a:\mu<8\). Write a contextual conclusion if \(\mu\) is the true mean charging time, in hours, for a specified device model.
  2. A two-sided test comparing two population means has \(p=0.30\) and \(\alpha=0.10\). What conclusion is appropriate, and why would “the population means are equal” overstate the result?
  3. Rewrite “We accept the null hypothesis that the true mean is 40 minutes” for a test with \(p>\alpha\) and \(H_a:\mu>40\).
  4. If a test fails to reject \(H_0:\mu=12\) against \(H_a:\mu\ne12\), what does the result say about evidence for a difference? What does it not prove?
  5. A p-value is \(0.07\) and the preselected significance level is \(0.05\). Explain how the conclusion would change if the significance level had instead been \(0.10\), and identify which value changes and which does not.