Tutorials › AP Statistics › Reject H0 Wording That Earns Full Credit

P-values and mean-inference conclusions · Tutorial 748 of 1000

Reject H0 Wording That Earns Full Credit

Practice turning a rejected null hypothesis into a clear, correctly qualified conclusion about a population mean or difference in means.

Intermediate 9 min read

What You'll Learn

  • Use the p-value and significance level to state when to reject the null hypothesis.
  • Connect a rejection directly to the direction or difference named in the alternative hypothesis.
  • Write a contextual conclusion about the population parameter, not just the sample result.
  • Explain a p-value conditionally on the null hypothesis and in the direction of the alternative.
  • Avoid claiming that rejection proves the alternative or establishes practical importance.
  • Adapt the conclusion wording to one-sided and two-sided tests of means.

A Rejection Needs a Complete Conclusion

In “Strength of Evidence on a Continuum,” you learned that a smaller p-value generally indicates stronger evidence against \(H_0\) for comparable tests. When a p-value is at or below the significance level \(\alpha\), the decision rule is to reject \(H_0\). The next step is to explain what that decision means for the population in the setting of the question.

A complete rejection conclusion does more than announce “reject.” It links the decision to the alternative hypothesis, describes the p-value as evidence against the null claim, and states the result in context. For a test about a mean, the conclusion concerns the true population mean, not simply the sample mean. For a comparison, it concerns the true difference between population means.

Conclusion pattern: Since the p-value is at or below \(\alpha\), reject \(H_0\). The p-value indicates that, if the null claim were true, results at least as extreme as those observed in the direction specified by \(H_a\) would be unusual. Therefore, the data provide convincing evidence in favor of \(H_a\): state its claim about the population parameter in the context of the problem.

The final sentence should match the alternative exactly. If \(H_a\) says a mean is less than a target, conclude there is convincing evidence that the population mean is less than that target. If \(H_a\) says two means differ, conclude there is convincing evidence that the population means differ. Do not turn a two-sided result into a claim about which mean is larger unless the hypotheses and observed data support that direction.

What Each Part of the Conclusion Does

A strong response usually has three connected parts. First, state the decision using the comparison between the p-value and \(\alpha\). Second, describe what the p-value says under \(H_0\). Third, explain what the evidence supports in the problem’s context, using the claim in \(H_a\).

1
Make the decision.
Compare the p-value with the preselected significance level. If \(p\le\alpha\), write “reject \(H_0\).”
2
Connect the p-value to the null.
Explain that the p-value is the probability, assuming \(H_0\) is true, of getting a test statistic at least as extreme as the observed one in the direction or directions specified by \(H_a\).
3
State the evidence in context.
Say that the data provide convincing evidence for the alternative claim about the true population mean or difference in means. Include the population and the variable, with units where useful.

The p-value explanation and the final evidence statement are related, but they do different jobs. The first describes how unusual the data would be if the null claim were true. The second answers the research question: what does the test provide evidence for? Including both makes the reasoning clear.

A conclusion does not need to repeat every calculation from the “Do” step. It should, however, make the result understandable without relying on a bare statement such as “the result is significant.” In particular, name the population mean or difference being tested and make the direction of the conclusion unmistakable.

Worked Example: Testing a Package-Fill Target

Worked Example: Testing a Package-Fill Target

A fictional packaging facility takes a random sample of 40 packages from a large production lot. The sample mean fill is \(498.5\) grams, and the sample standard deviation is \(4\) grams. The facility wants to know whether the true mean fill in the lot is below the labeled target of \(500\) grams. Assume the sample is less than 10% of the relevant production population and the sample data show no severe skewness or extreme outliers. Use \(\alpha=0.05\).

1
State.
Let \(\mu\) be the true mean fill, in grams, of packages in this production lot. The hypotheses are \(H_0:\mu=500\) grams and \(H_a:\mu<500\) grams.
2
Plan and check conditions.
Use a one-sample t test for a population mean. The packages were randomly sampled, so the random condition is met. The sample is less than 10% of the production lot, supporting independence under the 10% condition. With no severe skewness or extreme outliers, using a t procedure is reasonable. The left-sided alternative matches the question about underfilling.
3
Do.
With \(n=40\), the degrees of freedom are \(39\). The test statistic is
\(t=\dfrac{\bar{x}-500}{s/\sqrt{n}}=\dfrac{498.5-500}{4/\sqrt{40}}=\dfrac{-1.5}{0.63246}\approx-2.372.\)
The left-tail probability is \(P(T\le-2.372)\approx0.01137\), rounded, for \(39\) degrees of freedom. Thus the p-value is approximately \(0.01137\).
4
Conclude in context.
Because \(0.01137<0.05\), reject \(H_0\). If the true mean package fill were \(500\) grams, the probability of getting a t statistic of \(-2.372\) or lower would be about \(0.01137\). This is a small probability under \(H_0\), so the data provide convincing evidence that the true mean fill of packages in this production lot is less than \(500\) grams.

Notice that the conclusion names the true mean fill, not just the sample mean of \(498.5\) grams. It also follows the direction in \(H_a\): the evidence supports a mean below the target. The test does not establish how far below \(500\) grams the true mean is.

Match the Conclusion to the Alternative

The alternative hypothesis determines what a rejection supports. Keep the order and direction of the parameter in \(H_a\) visible as you write. This is especially important for a two-sample test, where reversing the order of subtraction reverses the meaning of a positive or negative difference.

Alternative hypothesisConclusion after rejecting \(H_0\)
\(H_a:\mu<\mu_0\)Convincing evidence that the true population mean is less than \(\mu_0\).
\(H_a:\mu>\mu_0\)Convincing evidence that the true population mean is greater than \(\mu_0\).
\(H_a:\mu\ne\mu_0\)Convincing evidence that the true population mean differs from \(\mu_0\).
\(H_a:\mu_1-\mu_2>0\)Convincing evidence that the true mean for population 1 is greater than the true mean for population 2.
\(H_a:\mu_1-\mu_2\ne0\)Convincing evidence that the two true population means differ.

For a two-sided alternative, “differs” is the safe conclusion even if the observed sample means have a particular order. You can mention that the sample mean in one group was higher, but do not present that observed order as the direction established by a two-sided alternative. If the question asks for the direction of the estimated difference, identify it as a sample result.

Worked Example: Evidence That a Mean Exceeds a Benchmark

A fictional transit agency randomly samples 25 days from a large set of operating days to study the mean number of minutes a particular route runs late. The sample mean is \(42.4\) minutes and the sample standard deviation is \(5\) minutes. The agency tests whether the true mean delay exceeds a \(40\)-minute benchmark. The sample is less than 10% of the relevant days, and the data have no severe skewness or extreme outliers. A one-sample t test reports \(p=0.012\), rounded to three decimal places. Use \(\alpha=0.05\).

State: Let \(\mu\) be the true mean delay, in minutes, for this route on the operating days represented by the study. The hypotheses are \(H_0:\mu=40\) minutes and \(H_a:\mu>40\) minutes.

Plan: A one-sample t test is appropriate. The days were randomly sampled; the sample is less than 10% of the population of relevant operating days; and the data show no severe skewness or extreme outliers. The right-sided alternative matches the question about delays exceeding the benchmark.

Do: The reported p-value, \(0.012\), is less than \(\alpha=0.05\), so reject \(H_0\). If the true mean delay were \(40\) minutes, the probability of obtaining a t statistic at least as large as the observed one would be about \(0.012\).

Conclude: The data provide convincing evidence that the true mean delay on this route is greater than \(40\) minutes. This conclusion is about the route’s population mean delay, not a claim that every day has a delay greater than \(40\) minutes.

Worked Example: A Two-Sided Test Comparing Mean Times

A fictional school takes independent random samples of students enrolled in two different study programs and compares their weekly study time. Let \(\mu_1\) and \(\mu_2\) be the true mean weekly study times, in hours, for students in programs 1 and 2 in the school. The sample means are \(18.6\) and \(21.0\) hours, respectively. A valid two-sample t test of \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\) reports \(p=0.008\). Both sample sizes are less than 10% of their respective populations, the groups are independent, and neither sample shows severe skewness or extreme outliers. Use \(\alpha=0.01\).

State: The hypotheses test whether the true mean weekly study times for the two program populations differ.

Plan: Use an unpooled two-sample t test; the problem states that the test is valid. The groups come from independent random samples, and each sample meets the 10% condition. The two-sided alternative is appropriate because the question asks whether the means differ in either direction.

Do: Since \(0.008<0.01\), reject \(H_0\). Assuming the two population means are equal, the probability of obtaining a test statistic at least as extreme as the observed one, in either direction, is \(0.008\).

Conclude: The data provide convincing evidence that the true mean weekly study times differ between students in the two programs at this school. The sample mean is lower in program 1 than in program 2, but the two-sided test’s conclusion is evidence of a difference, not proof that program 1 has the lower population mean. These samples also do not establish that enrollment in either program causes a change in study time.

Common Mistakes and AP Exam Tips

  • Stopping at “reject \(H_0\)”: That states the decision but not what the evidence means. Add a sentence naming the population parameter and the claim in \(H_a\).
  • Writing “accept \(H_a\)” or “prove \(H_a\)”: A test provides evidence, not proof. Prefer “the data provide convincing evidence in favor of \(H_a\)” followed by the contextual claim.
  • Describing the sample instead of the population: “The sample mean is below the target” is a description of the data. The inferential conclusion should say what the evidence suggests about the true population mean.
  • Ignoring the direction of \(H_a\): A left-tailed alternative supports a “less than” conclusion; a right-tailed alternative supports a “greater than” conclusion. A two-sided alternative supports a “differs” conclusion, not a directional claim by itself.
  • Calling the p-value the probability that \(H_0\) is true: The p-value is calculated assuming \(H_0\) is true. Describe the chance of results at least as extreme as those observed under that assumption.
  • Claiming practical importance from statistical significance alone: Rejecting \(H_0\) does not show that the difference is large or important. Discuss the estimated difference and its units, and use an appropriate confidence interval when available.
  • Forgetting the study design: A conclusion should not claim cause and effect from random sampling alone. Random assignment in an experiment can support a causal interpretation; a random sample supports generalization to its target population.

A concise, full-credit conclusion can be written in two or three sentences: “Since the p-value is ___ and \(\alpha\) is ___, reject \(H_0\). If the null claim were true, the probability of a test statistic at least as extreme as the observed one in the direction(s) of \(H_a\) would be ___. The data provide convincing evidence that [the contextual claim in \(H_a\)].” Fill in the final statement with the correct population and direction, rather than leaving it as a symbol-only conclusion.

Key takeaway: After rejecting \(H_0\), connect the p-value to the null claim and state that the data provide convincing evidence for the alternative hypothesis in context. Use the direction in \(H_a\), describe a population parameter, and avoid saying the test proves the alternative.

Check Your Understanding

For each question, focus on the decision, the meaning of the p-value under \(H_0\), and the contextual claim supported by \(H_a\).

  1. A one-sample test has \(p=0.018\) and \(\alpha=0.05\), with \(H_a:\mu<75\). Write a complete conclusion in context if \(\mu\) is the true mean battery life, in hours, for a specified model.
  2. A test comparing two population means has \(p=0.006\), \(\alpha=0.01\), and \(H_a:\mu_1-\mu_2\ne0\). What should the contextual conclusion say, and what directional claim should it avoid?
  3. Why is “reject \(H_0\), so the sample mean is significantly below the target” not a complete contextual conclusion?
  4. Rewrite “There is a 2% chance that the null hypothesis is true” for a test whose p-value is \(0.02\).
  5. A test rejects \(H_0\) at \(\alpha=0.05\). Does that alone show that the mean difference is practically important? Explain what other information could help.