Comparing Two Treatments with a Two-Sample t Test
When two treatments are given to different people, their recovery times form two independent groups if there is no pairing between the people in one group and those in the other. A two-sample t test can help assess whether the groups provide convincing evidence of different population mean recovery times.
In “The Two-Sample t Test Statistic” and “Finding the P-Value for a Two-Sample t Test,” you learned how to calculate the statistic and choose the appropriate tail area. Here, we put those pieces together in a full test. The question is whether mean recovery times differ, so the alternative is two-sided.
The order of subtraction must stay consistent. The statistic uses \(\bar{x}_1-\bar{x}_2\), so a negative statistic means the first treatment’s sample mean recovery time is lower than the second treatment’s. A two-sided alternative still considers results at least as far from zero in either direction.
Plan the Test Before Calculating
A complete response does more than report a p-value. First define the population means and hypotheses. Then explain why the study design and data support the two-sample t procedure. For a small sample, consider the shape of each group’s data separately; the distribution of paired differences is not relevant because the treatments are given to different people.
- Random design: The data should come from appropriate random samples or a randomized experiment.
- Independent groups and observations: Each person receives one treatment, and people’s recovery times can reasonably be treated as independent. There must not be a meaningful pairing between the groups.
- 10% condition: When sampling without replacement from a finite population, each sample should be less than 10% of its population. In a randomized experiment, check that recruiting a small fraction of the eligible population is reasonable when this applies.
- Nearly Normal condition: For small samples, each group’s distribution should be roughly symmetric with no strong outliers. Larger samples can better withstand departures from Normality.
As discussed in “Defining Both Population Means in Context,” define each parameter using the population, the quantitative variable, and its units. Also distinguish what the design permits you to say: random assignment can support a cause-and-effect conclusion about the treatments, while random sampling supports generalization to the population sampled.
Worked Examples
Worked Example: Comparing Mean Recovery Times
Suppose a hypothetical randomized trial compares two treatments for a particular injury. Twenty patients are randomly selected from a large population of eligible patients and randomly assigned, 10 to each treatment. Recovery time is measured in days. The Treatment 1 group has \(\bar{x}_1=12\) days and \(s_1=\sqrt{20}\) days; the Treatment 2 group has \(\bar{x}_2=17\) days and \(s_2=\sqrt{20}\) days. Plots of the recovery times show no strong skewness or outliers.
State: Let \(\mu_1\) be the true mean recovery time for eligible patients receiving Treatment 1, and let \(\mu_2\) be the true mean recovery time for eligible patients receiving Treatment 2. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\). The question asks whether the mean times differ, not whether one treatment has a shorter mean time.
Plan: Recovery time is quantitative. Each patient receives only one treatment, so the groups are independent rather than paired. Patients were randomly selected and randomly assigned. The 20 selected patients are less than 10% of the large eligible population, and the group plots show no strong skewness or outliers. These features support an unpooled two-sample t test. Random assignment supports a cause-and-effect conclusion about treatment for the study participants; random selection supports generalizing to the eligible population.
Do: Calculate the estimated standard error by keeping the two variance contributions separate:
The observed difference is \(12-17=-5\) days, so the test statistic is:
For these equal sample sizes and equal sample variances, the Welch degrees of freedom are 18. The calculation is:
Because the alternative is two-sided, find the area at least as far from zero as \(-2.5\) in both tails. Using the calculator’s t-distribution function:
This p-value means that if the population mean recovery times were equal, the probability of obtaining a two-sample t statistic with absolute value at least \(2.5\) is about \(0.02231\).
Conclude: At the 5% significance level, reject \(H_0\), because \(0.02231<0.05\). The trial provides convincing evidence that the true mean recovery times differ between the two treatments. Treatment 1’s sample mean is lower, but the conclusion is about a difference in means; it does not establish that every patient recovers faster with Treatment 1.
Worked Example: A Difference That Is Not Statistically Significant
In another hypothetical trial, different patients are randomly assigned to one of two treatments for a minor injury. Recovery times, in days, are recorded. Each group has \(n=6\). Treatment 1 has \(\bar{x}_1=8\) days and \(s_1=\sqrt{3}\) days; Treatment 2 has \(\bar{x}_2=8\) days and \(s_2=\sqrt{3}\) days. The plots show roughly symmetric distributions with no outliers.
State: Let \(\mu_1\) and \(\mu_2\) be the true mean recovery times for patients receiving Treatments 1 and 2. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\).
Plan: Recovery time is quantitative, patients are in only one group, and the groups have no matching or repeated measurements. Assume random assignment, and that the recruited patients are less than 10% of the eligible population. Since both sample sizes are small, the stated plots are important: they show no strong skewness or outliers. These conditions support a two-sample t test.
Do: Each variance contribution is \(s_i^2/n_i=3/6=0.5\), so the standard error is:
The observed difference and statistic are:
The Welch degrees of freedom are:
For a two-sided test with \(t=0\), the probability of a statistic at least as far from zero as the observed statistic is 1. Thus \(p=1.0000\).
Conclude: At the 5% significance level, fail to reject \(H_0\). These data do not provide convincing evidence that the population mean recovery times differ. The equal sample means do not prove that the true means are equal; they simply provide no observed difference to support the alternative.
Worked Example: Checking the Direction of a Two-Sided Result
A hypothetical rehabilitation center compares two recovery programs using independent groups of six patients. Recovery time is measured in days. Program A has \(\bar{x}_1=14\) and \(s_1=\sqrt{3}\); Program B has \(\bar{x}_2=12\) and \(s_2=\sqrt{3}\). Assume patients were randomly assigned, the sample is less than 10% of the eligible population, and plots show roughly symmetric distributions without outliers.
State: Let \(\mu_1\) and \(\mu_2\) be the true mean recovery times for patients receiving Programs A and B, respectively. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\).
Plan: The groups contain different patients, so they are independent, not paired. The response is quantitative, and the random assignment, 10% comparison, and plots support the two-sample t test.
Do: The standard error is \(\sqrt{3/6+3/6}=1\) day. Therefore:
The two-sided p-value is \(2\operatorname{tcdf}(2,1\text{E}99,10)\approx0.0734\), rounded. Although the statistic is positive, the alternative is two-sided, so the p-value includes equally extreme results in both directions.
Conclude: At the 5% significance level, fail to reject \(H_0\), since \(0.0734>0.05\). The sample results do not provide convincing evidence that the true mean recovery times differ between Programs A and B. The positive sample difference indicates that Program A’s sample mean was longer, but it is not strong enough evidence at this significance level to conclude that the population means differ.
What Makes the Conclusion Complete?
For a test, compare the p-value with the stated significance level \(\alpha\). If \(p\le\alpha\), reject \(H_0\); if \(p>\alpha\), fail to reject \(H_0\). Then answer the original question in context. As in “Finding the P-Value for a Two-Sample t Test,” do not interpret the p-value as the probability that the null hypothesis is true.
Keep the statistical conclusion aligned with the alternative. For \(H_a:\mu_1-\mu_2\ne0\), a significant result supports a difference, but a two-sided test by itself does not establish a specific direction. You can describe the observed direction by referring to \(\bar{x}_1-\bar{x}_2\), while making clear that this is the sample result.
The design also limits the wording. In a randomized experiment, a significant difference can support a claim that the treatments caused different mean outcomes for the study population. Generalizing that claim to a wider population depends on how participants were selected. A randomized sample without random assignment supports generalization but does not, by itself, establish that a treatment caused the difference.
Common Mistakes and AP Exam Tips
- Using a paired test for separate patients: If each patient receives only one treatment and there is no matching, use a two-sample test. A paired t test requires one meaningful difference per pair.
- Choosing a one-sided alternative after seeing the data: If the research question asks whether the means differ, use \(H_a:\mu_1-\mu_2\ne0\), even when one sample mean is larger.
- Reversing the group order: Define \(\mu_1-\mu_2\) first, then use that same order for the sample means and statistic. Changing the order changes the sign of \(t\), but not the two-sided p-value.
- Skipping conditions: A complete plan explains the random design, independence, 10% condition when relevant, and distribution shape, especially for small samples.
- Using pooled degrees of freedom: The standard unpooled two-sample t test uses Welch degrees of freedom from the calculator, not \(n_1+n_2-2\).
- Saying “accept the null”: When the p-value is not small, say “fail to reject \(H_0\)” and explain that the data do not provide convincing evidence of a difference. Do not claim the means are proven equal.
- Leaving out context: Full-credit communication identifies the recovery-time means, names the treatments, states the significance level, and gives the conclusion in the setting’s units and context.
Check Your Understanding
Use the two-sample t test ideas from the examples to answer these questions.
- Two independent treatment groups have sample means of 9 and 13 days. If \(\mu_1\) and \(\mu_2\) follow that same group order, what is the sign of \(\bar{x}_1-\bar{x}_2\)?
- A study asks whether mean recovery times differ. Should the alternative be one-sided or two-sided? Write it using \(\mu_1\) and \(\mu_2\).
- Why would two groups of different patients who each receive only one treatment generally call for a two-sample test rather than a paired t test?
- If a test gives \(p=0.08\) at \(\alpha=0.05\), what decision should you make, and what should you avoid claiming?
- In a randomized treatment experiment, what does random assignment support, and what does random sampling support?