Tutorials › AP Statistics › Writing a Conclusion for a Two-Sample t Test

Two-sample t hypothesis tests · Tutorial 733 of 1000

Writing a Conclusion for a Two-Sample t Test

Connect the p-value and significance level to a clear, contextual conclusion about two population means.

Intermediate 9 min read

What You'll Learn

  • Compare a two-sample t test p-value with the stated significance level.
  • Write a conclusion that names the populations, measured variable, and direction of the difference.
  • Distinguish the sample mean difference from the population mean difference.
  • Explain what “fail to reject” does—and does not—mean in context.
  • Match a one-sided conclusion to the direction in the alternative hypothesis.
  • Avoid claiming that a small p-value measures the size or importance of a difference.

Turning a Test Result into a Conclusion

In “Interpreting the Test Statistic for Two Means,” we saw how the t statistic describes the observed difference in sample means relative to its estimated standard error. In “Finding the P-Value for a Two-Sample t Test,” we used the alternative hypothesis and the t distribution to find the p-value. Now we connect that p-value to a conclusion about the population means.

A complete conclusion does more than say “reject” or “fail to reject.” It states the decision relative to the significance level, then explains what the evidence indicates about the population mean difference in the setting of the question. Keep the group order consistent: if the parameter is \(\mu_1-\mu_2\), describe population 1’s mean relative to population 2’s mean.

Definition: The significance level, \(\alpha\), is the threshold chosen for deciding whether a p-value provides convincing evidence against the null hypothesis. If the p-value is less than or equal to \(\alpha\), reject \(H_0\). If the p-value is greater than \(\alpha\), fail to reject \(H_0\).

A Reliable Conclusion Pattern

For a two-sample t test of \(H_0:\mu_1-\mu_2=0\), the null hypothesis says the population means are equal. The alternative hypothesis describes the kind of difference the test is looking for: a difference in either direction, a greater mean for population 1, or a smaller mean for population 1. As covered in “Choosing One-Sided or Two-Sided for Two Means,” the conclusion must follow that alternative.

Use the p-value and \(\alpha\) to make the decision. A p-value at or below \(\alpha\) leads to rejecting \(H_0\); a p-value above \(\alpha\) leads to failing to reject \(H_0\). Then give the conclusion in context. When you reject \(H_0\), state that there is convincing evidence for the difference described by \(H_a\). When you fail to reject \(H_0\), state that there is not convincing evidence for that difference. Neither decision proves a claim about the population means.

1
Compare.
Compare the p-value with the stated \(\alpha\). A p-value equal to \(\alpha\) is on the rejection side of the decision rule.
2
Decide.
Reject \(H_0\) if \(p\leq\alpha\); otherwise, fail to reject \(H_0\).
3
Conclude in context.
Name the populations and quantitative variable, and describe the mean difference in the direction specified by \(H_a\).

The p-value is calculated assuming \(H_0\) is true. It is the probability of getting a test statistic at least as extreme as the observed one in the direction or directions specified by \(H_a\). It is not the probability that \(H_0\) is true, and it does not tell us how large or practically important the population difference is.

The sample mean difference, \(\bar{x}_1-\bar{x}_2\), does have the variable’s original units and shows the direction of the observed difference. It is useful to include when it helps explain the result, but do not confuse it with the unknown population difference, \(\mu_1-\mu_2\). A test conclusion is about the population parameters, not just the two sample means.

Conclusion template: Because the p-value is [less than or equal to / greater than] \(\alpha=[\text{value}]\), [reject \(H_0\) / fail to reject \(H_0\)]. There [is / is not] convincing evidence that [state the difference in population means described by \(H_a\)].

Worked Examples

Worked Example: Evidence of a Difference in Mean Activity Time

A hypothetical survey uses independent random samples of 30 students from each of two large schools. Students report the number of minutes they spend on a weekly activity. School 1’s sample mean is 41.2 minutes, and School 2’s is 37.2 minutes. Both sample standard deviations are 6 minutes. The question is whether the population means differ. Use \(\alpha=0.05\), and suppose the two-sample t test gives a two-sided p-value of 0.0124.

State: Let \(\mu_1\) and \(\mu_2\) be the true mean weekly activity times for students at Schools 1 and 2, respectively. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\).

Plan: Use an unpooled two-sample t test. The samples are independent random samples, and the schools have large student populations, so each sample is less than 10% of its population. The groups are independent because students are sampled separately and are not paired. Both sample sizes are 30, so the t procedure is reasonably robust to departures from Normality, provided there are no severe outliers or extreme skewness. Assume inspection of the sample distributions shows no severe outliers or extreme skewness. These conditions support the test.

Do: The observed difference in the specified order is \(41.2-37.2=4.0\) minutes. The standard error is

$$ SE_{\bar{x}_1-\bar{x}_2} =\sqrt{\frac{6^2}{30}+\frac{6^2}{30}} =\sqrt{1.2+1.2} =\sqrt{2.4} \approx 1.5492\text{ minutes} $$

The test statistic is \(t=4.0/1.5492\approx2.582\). Using the unpooled two-sample t procedure, the calculator gives approximately 58 degrees of freedom and a two-sided p-value of 0.0124, rounded to four decimal places. Since \(0.0124<0.05\), reject \(H_0\).

Conclude: There is convincing evidence that the true mean weekly activity times differ between students at the two schools. The observed difference is positive: the sample mean for School 1 is 4.0 minutes greater than the sample mean for School 2. The test supports a difference between the population means, but the observed 4.0-minute difference is not an exact measurement of the population difference.

Worked Example: Not Enough Evidence to Establish a Difference

A hypothetical environmental project compares the weekly hours of sunlight received by randomly selected garden plots in two neighborhoods. The groups are independent, with 20 plots in each sample. Neighborhood 1 has a sample mean of 6.5 hours and a standard deviation of 5 hours; Neighborhood 2 has a sample mean of 5.5 hours and a standard deviation of 5 hours. The research question asks whether the population means differ. Use \(\alpha=0.05\).

Let \(\mu_1\) and \(\mu_2\) be the true mean weekly sunlight hours for garden plots in Neighborhoods 1 and 2, respectively. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\).

The samples are random and independent, and each is less than 10% of its neighborhood’s plots. With 20 observations per group, inspect each group’s distribution for strong skewness or outliers; assume neither sample has a feature that would make the t procedure inappropriate. The conditions support an unpooled two-sample t test.

The observed difference is \(6.5-5.5=1.0\) hour. The standard error and test statistic are

$$ SE_{\bar{x}_1-\bar{x}_2} =\sqrt{\frac{5^2}{20}+\frac{5^2}{20}} =\sqrt{1.25+1.25} =\sqrt{2.5} \approx1.5811\text{ hours} $$
$$ t=\frac{1.0}{1.5811}\approx0.6325 $$

The calculator’s unpooled test gives 38 degrees of freedom and a two-sided p-value of approximately 0.5313, rounded to four decimal places. Since \(0.5313>0.05\), fail to reject \(H_0\).

There is not convincing evidence that the true mean weekly sunlight hours differ between garden plots in the two neighborhoods. This does not establish that the population means are equal. The test simply did not find sufficiently strong evidence of a difference at the 0.05 significance level.

Worked Example: A One-Sided Conclusion About Which Mean Is Greater

In a hypothetical experiment, 10 comparable devices are randomly assigned to each of two battery settings. Researchers measure operating time in hours. Setting 1 has a sample mean of 13.5777 hours, and Setting 2 has a sample mean of 10.0000 hours. Both sample standard deviations are 4 hours. Researchers want to know whether Setting 1 produces a greater mean operating time. Use \(\alpha=0.05\).

Let \(\mu_1\) and \(\mu_2\) be the true mean operating times for devices using Settings 1 and 2, respectively. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2>0\).

Random assignment makes the treatment groups suitable for comparing the settings. Each device receives only one setting, so the groups are independent. Because each group has only 10 devices, the t procedure requires checking the response distributions for strong skewness and outliers. Assume plots show no strong skewness or outliers. These features, along with the randomized design, support an unpooled two-sample t test.

The observed difference is \(13.5777-10.0000=3.5777\) hours. The standard error and statistic are

$$ SE_{\bar{x}_1-\bar{x}_2} =\sqrt{\frac{4^2}{10}+\frac{4^2}{10}} =\sqrt{1.6+1.6} =\sqrt{3.2} \approx1.7889\text{ hours} $$
$$ t=\frac{3.5777}{1.7889}\approx2.000 $$

The calculator’s unpooled test gives 18 degrees of freedom. Because \(H_a\) is right-tailed, use the area to the right of \(t=2.000\), not a two-sided area. The p-value is approximately 0.0304, rounded to four decimal places. Since \(0.0304<0.05\), reject \(H_0\).

There is convincing evidence that, for the devices in this experiment, Setting 1 produces a greater mean operating time than Setting 2. This conclusion follows the direction in the alternative hypothesis. It does not claim that every device using Setting 1 lasts longer than every device using Setting 2.

Common Mistakes and AP Exam Tips

A good conclusion is concise, but it includes the decision and its meaning in context. Avoid turning the p-value into a claim it does not support.

  • Writing only “reject” or “fail to reject”: Add a statement about convincing evidence for the population mean difference in context.
  • Claiming that the null hypothesis is true: “Fail to reject” does not mean “accept \(H_0\)” or prove the population means are equal. Say that there is not convincing evidence for the difference described by \(H_a\).
  • Reversing the groups: If the test concerns \(\mu_1-\mu_2\), describe population 1 relative to population 2. Check that the direction in your conclusion matches both the subtraction order and \(H_a\).
  • Using the wrong tail for a one-sided test: A right-tailed alternative asks whether \(\mu_1\) is greater than \(\mu_2\); a left-tailed alternative asks whether it is smaller. Do not double a one-tail p-value unless the test is two-sided.
  • Calling the p-value the chance that the null is true: State that the p-value is calculated assuming \(H_0\) is true. It measures how unusual the observed test statistic, or one more extreme, would be under that assumption.
  • Confusing statistical evidence with size or importance: A small p-value is evidence against the null, not a measure of the difference’s size. If useful, report the sample mean difference with units, while making clear that the test concerns the population means.
  • Ignoring the stated significance level: The decision depends on comparing the p-value with the specified \(\alpha\). Do not use a memorized cutoff when the problem gives a different level.
Key takeaway: Compare the p-value with \(\alpha\), make the reject-or-fail-to-reject decision, and then describe the evidence for the population mean difference in context and in the direction of \(H_a\). A test result does not prove the null hypothesis or measure the practical size of a difference.

Check Your Understanding

For each question, connect the decision to the stated alternative and population means.

  1. A two-sided test has p-value 0.08 and \(\alpha=0.10\). State the decision and write a contextual conclusion about the population means.
  2. A right-tailed test of \(H_a:\mu_1-\mu_2>0\) has p-value 0.12 and \(\alpha=0.05\). What decision should be made, and what should the conclusion say?
  3. Why is “there is a 3% chance that the null hypothesis is true” not a correct interpretation of a p-value of 0.03?
  4. A test rejects \(H_0\), and \(\bar{x}_1-\bar{x}_2=-2.4\) kilograms. What direction did the sample difference have, and what else must you check before writing the conclusion?
  5. Explain why failing to reject \(H_0:\mu_1-\mu_2=0\) does not demonstrate that the population means are equal.