Tutorials › AP Statistics › Worked Example: One-Sided Test for Two Means

Two-sample t hypothesis tests · Tutorial 729 of 1000

Worked Example: One-Sided Test for Two Means

Work through right-tailed two-sample t tests to assess whether one population mean is greater than another.

Intermediate 9 min read

What You'll Learn

  • Define the population mean difference to match a claim that one group’s mean exceeds another’s.
  • Write null and right-tailed alternative hypotheses using a consistent group order.
  • Check the conditions for an unpooled two-sample t test in context.
  • Calculate the standard error, test statistic, Welch degrees of freedom, and right-tail p-value.
  • Interpret significant and nonsignificant results without overstating what the data show.

Testing a Directional Claim About Two Means

A question about whether one group’s population mean is greater than another’s calls for a right-tailed test when the difference is defined in that same order. In this tutorial, group 1 is the group whose mean is claimed to be greater. The test statistic will use \(\bar{x}_1-\bar{x}_2\), so positive values point in the direction of the claim.

As in “Choosing One-Sided or Two-Sided for Two Means,” the alternative hypothesis should reflect the research question and be chosen before examining the results. A right-tailed test is not a way to make a result look more persuasive after noticing which sample mean is larger. The earlier tutorials “The Two-Sample t Test Statistic” and “Finding the P-Value for a Two-Sample t Test” explain how to calculate the statistic and select a tail area; here, we use those ideas in complete right-tailed examples.

Definition: Let \(\mu_1\) be the true population mean for group 1 and \(\mu_2\) the true population mean for group 2, for the same quantitative variable. To test whether group 1’s population mean exceeds group 2’s, use \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2>0\). The null states no difference; the alternative states the direction of the claim.

The order matters throughout. With these hypotheses, a positive test statistic is in the direction of the alternative, while a negative statistic is in the opposite direction. The p-value is the area to the right of the observed statistic under the null distribution. It is not a two-tail area, even if the observed difference happens to point away from the claim.

Right-Tail Test: The Steps

For an unpooled two-sample t test, the observed difference in sample means is compared with the null difference of zero. The standard error estimates the variability in that difference, and the Welch degrees of freedom determine which t distribution is used for the p-value.

$$ SE=\sqrt{\frac{s_1^2}{n_1}+\frac{s_2^2}{n_2}}, \qquad t=\frac{(\bar{x}_1-\bar{x}_2)-0}{SE}. $$

For a right-tailed alternative, the p-value is the probability, assuming the null hypothesis is true, of getting a test statistic at least as large as the observed one. Use the statistic and Welch degrees of freedom from the test.

$$ p=\operatorname{tcdf}(t_{\text{obs}},1\text{E}99,df). $$

A small p-value means the observed result, or a result even more in the direction of the alternative, would be unusual if the population means were equal. It does not give the probability that the null hypothesis is true.

Conditions:
  • Random design: The data should come from appropriate random samples or a randomized experiment.
  • Independent groups and observations: Observations in the two groups are not paired, and observations within each group can reasonably be treated as independent.
  • 10% condition: For samples drawn without replacement, each sample should be less than 10% of its population.
  • Nearly Normal condition: With small samples, inspect each group’s distribution separately. Rough symmetry and no strong outliers support the procedure; larger samples can better withstand departures from Normality.

As in “Conditions for a Two-Sample t Interval,” describe the design, independence, 10% condition when relevant, and distribution shape in context. Do not assess the shape of differences: these are independent groups, not paired data. Also distinguish random assignment, which can support a cause-and-effect conclusion, from random sampling, which can support generalization to a population.

Worked Examples

Worked Example: Does a New Soil Mix Increase Mean Seedling Height?

Suppose a hypothetical greenhouse experiment compares seedlings grown in a new soil mix with seedlings grown in a standard mix. Twenty plots are randomly selected from a large set of eligible plots, then randomly assigned, 10 to each mix. After a set growing period, seedling height is measured in centimeters. The new-mix group has \(\bar{x}_1=48\) cm and \(s_1=\sqrt{20}\) cm; the standard-mix group has \(\bar{x}_2=43\) cm and \(s_2=\sqrt{20}\) cm. Plots of the heights show no strong skewness or outliers. Use \(\alpha=0.05\).

State: Let \(\mu_1\) be the true mean seedling height after the growing period for plots using the new mix, and \(\mu_2\) the true mean height for plots using the standard mix. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2>0\). The claim is specifically that the new mix produces a greater population mean height.

Plan: Height is quantitative. Different plots receive the two mixes, with no matching between groups, so the groups are independent. Plots were randomly sampled and randomly assigned. The 10 plots in each group are less than 10% of the large eligible set, and the plots can reasonably be treated as independent. Because the samples are small, the stated plots matter; they show no strong skewness or outliers in either group. These conditions support an unpooled two-sample t test. Random assignment supports a cause-and-effect conclusion about the mixes, and random sampling supports generalizing to the eligible plots in the greenhouse setting.

Do: The two variance contributions are \(s_1^2/n_1=20/10=2\) and \(s_2^2/n_2=20/10=2\). Thus:

$$ SE=\sqrt{\frac{20}{10}+\frac{20}{10}} =\sqrt{4}=2\text{ cm}. $$

The observed difference is \(48-43=5\) cm, in the direction of the claim. The test statistic is:

$$ t=\frac{(48-43)-0}{2}=2.5. $$

The Welch degrees of freedom are:

$$ df= \frac{(2+2)^2} {\frac{2^2}{10-1}+\frac{2^2}{10-1}} = \frac{16}{\frac{4}{9}+\frac{4}{9}} = \frac{16}{8/9} =18. $$

Because the alternative is right-tailed, use only the area to the right of \(t=2.5\):

$$ p=\operatorname{tcdf}(2.5,1\text{E}99,18) \approx 0.0112\text{ (rounded)}. $$

If the population mean heights were equal, the probability of obtaining a test statistic of \(2.5\) or greater in the direction of the claim would be about \(0.0112\).

Conclude: Since \(0.0112<0.05\), reject \(H_0\). The experiment provides convincing evidence that the true mean seedling height after the growing period is greater for plots using the new soil mix than for plots using the standard mix.

Worked Example: A Higher Sample Mean Is Not Enough

In a hypothetical randomized study, different students are assigned to one of two short practice routines, and their improvement on a balance task is measured in points. The research question, set before the study, is whether Routine 1 leads to a greater mean improvement than Routine 2. Each group has \(n=8\). Routine 1 has \(\bar{x}_1=31\) points and \(s_1=2\) points; Routine 2 has \(\bar{x}_2=30\) points and \(s_2=2\) points. The small-sample plots are roughly symmetric without outliers. Use \(\alpha=0.05\).

State: Let \(\mu_1\) be the true mean improvement for students using Routine 1 and \(\mu_2\) the true mean improvement for students using Routine 2. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2>0\).

Plan: Improvement is quantitative, and each student uses only one routine, so the groups are independent rather than paired. Suppose students were randomly selected from the target student population and randomly assigned to the routines. The 8 students in each sample are less than 10% of that population, and observations can reasonably be treated as independent. The roughly symmetric plots without outliers support the Nearly Normal condition for these small samples. A two-sample t test is appropriate.

Do: The standard error is:

$$ SE=\sqrt{\frac{2^2}{8}+\frac{2^2}{8}} =\sqrt{0.5+0.5}=1\text{ point}. $$

The observed difference is \(31-30=1\) point, so:

$$ t=\frac{(31-30)-0}{1}=1. $$

For the Welch degrees of freedom, each variance contribution is \(0.5\):

$$ df= \frac{(0.5+0.5)^2} {\frac{0.5^2}{7}+\frac{0.5^2}{7}} = \frac{1}{0.25/7+0.25/7} =14. $$

The right-tail p-value is:

$$ p=\operatorname{tcdf}(1,1\text{E}99,14) \approx 0.1671\text{ (rounded)}. $$

Assuming the population means are equal, a statistic of \(1\) or greater in the direction of Routine 1’s claimed advantage would occur with probability about \(0.1671\).

Conclude: Since \(0.1671>0.05\), fail to reject \(H_0\). The study does not provide convincing evidence that students using Routine 1 have a greater true mean improvement than students using Routine 2. Routine 1’s sample mean was higher, but that observed difference is not strong enough evidence for the directional claim.

Worked Example: The Sample Difference Points the Other Way

A hypothetical randomized experiment compares the time, in minutes, that students take to complete a design task using two software interfaces. The question, chosen before data are collected, is whether Interface A has a greater population mean completion time than Interface B. Each group contains 9 different students. Interface A has \(\bar{x}_1=24\) minutes and \(s_1=3\) minutes; Interface B has \(\bar{x}_2=26\) minutes and \(s_2=3\) minutes. The group plots show roughly symmetric distributions with no outliers. Use \(\alpha=0.05\).

State: Let \(\mu_1\) be the true mean completion time for students using Interface A and \(\mu_2\) the true mean completion time for students using Interface B. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2>0\). The alternative asks whether A’s mean time is greater, not whether the interfaces differ in either direction.

Plan: Completion time is quantitative. Each student uses one interface, so the groups are independent and unpaired. Suppose students were randomly selected from the target population and randomly assigned to an interface. Each sample is less than 10% of that population, and students’ completion times can reasonably be treated as independent. The small-sample plots show no strong skewness or outliers. These conditions support the unpooled two-sample t test.

Do: The standard error is:

$$ SE=\sqrt{\frac{3^2}{9}+\frac{3^2}{9}} =\sqrt{1+1} =\sqrt{2}\approx1.4142\text{ minutes}. $$

The observed difference is \(24-26=-2\) minutes. Therefore:

$$ t=\frac{(24-26)-0}{\sqrt{2}} \approx -1.4142. $$

The Welch degrees of freedom are:

$$ df= \frac{(1+1)^2} {\frac{1^2}{9-1}+\frac{1^2}{9-1}} = \frac{4}{1/8+1/8} =16. $$

Use the right-tail area even though the observed statistic is negative:

$$ p=\operatorname{tcdf}(-1.4142,1\text{E}99,16) \approx 0.9118\text{ (rounded)}. $$

If the means were equal, a test statistic at least as large as \(-1.4142\), in the direction of the claim, would be quite likely. This large p-value does not support the claim that Interface A has a greater mean completion time.

Conclude: Since \(0.9118>0.05\), fail to reject \(H_0\). The data do not provide convincing evidence that the true mean completion time is greater for students using Interface A. The sample mean for A is lower, but this right-tailed test is not a test of whether B’s mean is greater; that would require a different alternative specified before examining the data.

Common Mistakes and AP Exam Tips

A complete response makes the direction of the test visible from beginning to end. State which population is group 1, write the alternative using the same order, calculate \(\bar{x}_1-\bar{x}_2\), and use the matching right-tail area. A correct p-value with hypotheses that do not match the question is not a complete solution.

  • Choosing the direction after seeing the sample means: A one-sided alternative must come from the question or study plan, not from whichever sample mean turned out larger. Explain the research claim in context before calculating.
  • Using two tails: For \(H_a:\mu_1-\mu_2>0\), use the area to the right of the observed \(t\), not twice a tail area. The p-value is still a right-tail area when \(t\) is negative.
  • Reversing the group order: If group 1 is the group claimed to have the greater mean, keep that order in the parameters, sample means, and statistic. Reversing the subtraction reverses the sign and changes which tail matches the claim.
  • Concluding that a nonsignificant result proves equality: Say “fail to reject \(H_0\)” and explain that the data do not provide convincing evidence for the stated alternative. Do not say the population means are equal or that the null has been proven.
  • Ignoring the study design: State why the samples are independent and address the random design, 10% condition when applicable, and distribution shape for small samples. Random assignment and random sampling support different kinds of conclusions.
  • Reporting a direction stronger than the test supports: A significant right-tailed result supports evidence for \(\mu_1>\mu_2\). It does not show that every individual in group 1 has a larger value, or that the difference is practically important.
Key takeaway: For a right-tailed two-sample t test, define group 1 as the group whose greater mean is claimed, use \(H_a:\mu_1-\mu_2>0\), check the independent-sample conditions, and calculate the right-tail p-value using the observed statistic and Welch degrees of freedom. Conclude in context and do not overstate what the evidence shows.

Check Your Understanding

Use the right-tailed two-sample t test ideas from the examples to answer these questions.

  1. A study asks whether group A has a greater population mean than group B. If \(\mu_1\) and \(\mu_2\) refer to A and B in that order, write the null and alternative hypotheses.
  2. A right-tailed test has an observed statistic \(t=-0.8\). Should the p-value be a left-tail or right-tail area, and how should it be interpreted?
  3. Why should a researcher decide on a one-sided alternative before looking at the sample means?
  4. A test gives \(p=0.03\) with \(\alpha=0.05\). State the decision and conclusion structure without claiming that every individual in one group has a larger value.
  5. For two independent samples, which groups’ distributions should be examined for the Nearly Normal condition when both sample sizes are small?