Three Places a Two-Sample t Test Can Go Wrong
A two-sample t test can be calculated correctly and still answer the wrong question. Errors often begin before the calculator is used: the hypotheses may not match the research claim, the data may be treated as paired when they are not, or the p-value may be described as a probability that the null hypothesis is true.
In “Pooled Two-Sample t Test and Why AP Avoids It,” you saw that the usual AP method is the unpooled Welch two-sample t test. This tutorial focuses on checking whether the hypotheses, procedure, and interpretation fit the study. The key habit is to trace each part back to the same population comparison and study design.
Match the Hypotheses to the Question
For two independent groups, define \(\mu_1\) and \(\mu_2\) as the population means for the same quantitative variable, with group 1 and group 2 identified clearly. A test of no difference uses \(H_0:\mu_1-\mu_2=0\). The alternative states the difference the research question is asking about: a difference in either direction, or a difference in one specified direction.
A common error is to write a directional alternative based on which sample mean happens to be larger. The alternative should reflect the question or claim, not a direction selected after looking at the sample results. Another common error is to reverse the group order without reversing the sign in the alternative. If group 1 is expected to have a larger mean, use \(H_a:\mu_1-\mu_2>0\), not \(H_a:\mu_1-\mu_2<0\).
Worked Example: Fixing a Reversed Alternative
A fictional school randomly samples 12 students from each of two independent study programs and records the number of minutes each student spends on a weekly practice activity. The research question is whether students in Program A spend more time on average than students in Program B. The sample summaries are \(\bar{x}_1=53.266\) minutes and \(s_1=4\) minutes for Program A, and \(\bar{x}_2=50.000\) minutes and \(s_2=4\) minutes for Program B. Test at \(\alpha=0.05\).
Let \(\mu_1\) be the true mean weekly practice time for all students in Program A at this school, and let \(\mu_2\) be the true mean weekly practice time for all students in Program B at this school. The claim that Program A students spend more time corresponds to \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2>0\).
Use an unpooled two-sample t test because the samples are independent, not paired. The setting specifies random samples from the two programs. Assume each sample is less than 10% of its program’s student population, so the 10% condition is met. With 12 students in each group, check the distributions carefully; suppose plots show no strong skewness or outliers in either group. These conditions support the t procedure.
The sample difference is \(53.266-50.000=3.266\) minutes. The unpooled standard error and test statistic are:
The Welch degrees of freedom are 22 because the sample sizes and standard deviations are equal. For the right-tailed alternative, the p-value is the area to the right of \(t=2.000\) under a t distribution with 22 degrees of freedom. A calculator gives \(p\approx0.0290\), rounded.
Because \(0.0290<0.05\), reject \(H_0\). These data provide convincing evidence that the true mean weekly practice time is greater for Program A students than for Program B students at this school.
A reversed alternative, \(H_a:\mu_1-\mu_2<0\), would test the opposite claim. It would not become appropriate just because the sample difference had gone the other way; the research question sets the direction before the results are examined. The conclusion here is about a difference in means, not proof that attending one program causes students to practice more.
Do Not Use a Paired Formula Without Pairs
As explained in “Identifying Two Independent Samples,” independent groups contain different observational units with no design-based link between the groups. A paired t procedure instead requires one consistently defined difference for each genuine pair, as covered in “Common Mistakes With Paired t Procedures.” Matching two unrelated observations after data collection does not create pairs.
This distinction changes what information the procedure uses. For independent groups, the unpooled standard error is \(\sqrt{s_1^2/n_1+s_2^2/n_2}\). For paired data, the standard error is based on the standard deviation of the pairwise differences, \(s_d/\sqrt{n}\), where \(n\) counts pairs. If the study has no pairs, there is no meaningful \(s_d\) to substitute into the paired formula.
Worked Example: Independent Groups Are Not Paired
A fictional recreation center randomly samples 10 members who use a rowing machine and 10 different members who use a stationary bike. The response is minutes of exercise per visit. The sample summaries are \(\bar{x}_1=18\), \(s_1=4\) minutes for rowing, and \(\bar{x}_2=15\), \(s_2=3\) minutes for biking. The samples are independent; no member used both machines as part of a matched design.
A student tries to use a paired t procedure with \(n=10\) and \(s_d\). That is not valid: the study contains no defined pairwise differences, so \(s_d\) cannot be calculated from the information given. Pairing the first rowing member with the first biking member, for example, would be an arbitrary match with no basis in the design.
The appropriate standard error for the independent samples is
For a test of equal population means, the observed difference is \(18-15=3\) minutes, so
The Welch degrees of freedom are approximately
This example illustrates a procedure choice, not a complete claim about the recreation center’s population means: the research question and significance level would also be needed to state hypotheses and reach a test conclusion. The essential correction is that independent samples call for the unpooled two-sample standard error, not a paired standard error.
Say Exactly What the P-Value Means
A p-value is calculated under the assumption that the null hypothesis is true. It describes the probability of getting a test statistic at least as extreme as the observed statistic, in the direction specified by the alternative. The alternative determines which tail area—or areas—to count. “Finding the P-Value for a Two-Sample t Test” develops these tail calculations.
A p-value is not the probability that the null hypothesis is true. It is also not the probability that the observed sample difference happened “by chance.” Those statements confuse a probability about possible sample results, assuming a hypothesis, with a probability about the hypothesis itself.
Worked Example: Correcting a P-Value Interpretation
Return to the study-program test. Its test statistic was approximately \(t=2.000\), with 22 degrees of freedom, and its right-tailed p-value was approximately \(0.0290\). A student writes: “There is a 2.90% chance that the two programs have equal mean practice times.”
That statement is incorrect because the p-value does not give the probability that \(H_0\) is true. The correct interpretation is: Assuming the true mean weekly practice times are equal for the two programs, the probability of obtaining a sample difference favoring Program A at least as strongly as the observed difference—equivalently, a t statistic of 2.000 or greater—is about 0.0290.
The tail matters. If the research question had instead been whether the program means differ in either direction, the alternative would be two-sided, \(H_a:\mu_1-\mu_2\ne0\). For the same observed statistic and degrees of freedom, the two-sided p-value is approximately \(2(0.0290)=0.0580\), rounded. It is not the same question as the right-tailed test, so the p-value and test decision can differ at \(\alpha=0.05\).
The correct conclusion also uses the selected alternative and context. For the stated right-tailed test, \(p\approx0.0290\) is below \(0.05\), so there is convincing evidence that Program A’s population mean practice time is greater. The result does not establish a cause-and-effect relationship: the study sampled students from existing programs rather than randomly assigning them to programs.
Common Mistakes and AP Exam Tips
- Writing hypotheses about sample means: Use population parameters such as \(\mu_1\) and \(\mu_2\), not \(\bar{x}_1\) and \(\bar{x}_2\). A hypothesis is a claim about populations.
- Letting the sample dictate the alternative: Choose the direction from the research question before looking at the observed difference. The sample result is evidence for evaluating the alternative, not a reason to rewrite it.
- Changing group order but not the sign: If you switch which group is population 1, rewrite the difference and alternative consistently. A claim that Program A’s mean exceeds Program B’s mean is \(\mu_A-\mu_B>0\).
- Calling independent observations pairs: A shared sample size does not make two groups paired. Look for a genuine link, such as repeated measurements on the same person or matched units. Without that link, do not invent differences.
- Using the wrong standard error: For independent groups, keep the variance contributions separate as \(s_1^2/n_1+s_2^2/n_2\). A paired procedure instead uses the variability among actual pairwise differences.
- Misreading the p-value: Do not say it is the probability that the null hypothesis is true or that the result is due to chance. State what sample outcome would be at least as extreme, under \(H_0\), and follow the alternative’s tail.
- Making the conclusion too broad: A small p-value supports evidence about the population means in context. It does not, by itself, prove causation or guarantee that the difference is practically important.
A reliable final audit is to line up the parameter definitions, group order, hypotheses, test type, tail, and conclusion. Each should express the same comparison. For a full-credit response, identify the populations and quantitative variable, explain why the two-sample procedure fits the design, show the test statistic and p-value, and make the reject-or-fail-to-reject decision in context.
Check Your Understanding
For each question, explain how you would avoid the error in a two-sample t test.
- A research question asks whether Group B has a greater population mean than Group A. If you define \(\mu_1\) for Group A and \(\mu_2\) for Group B, what is the correct directional alternative?
- Two independent classes have equal sample sizes. Does that make their observations paired? What feature of the design would support pairing?
- Why can’t you use \(s_d/\sqrt{n}\) if a study has independent groups but no defined pairs?
- In a right-tailed test with p-value 0.04, what does the 0.04 describe, assuming the null hypothesis is true?
- Why does rejecting equal population means in a comparison of existing programs not, by itself, show that the program caused the difference?