Tutorials › AP Statistics › The Two-Sample t Test Statistic

Two-sample t hypothesis tests · Tutorial 724 of 1000

The Two-Sample t Test Statistic

Use the difference in sample means and the unpooled standard error to calculate a t statistic that measures how far the observed difference is from the null value of zero.

Intermediate 9 min read

What You'll Learn

  • Identify the observed difference in sample means and the null difference used in a two-sample t test.
  • Calculate the unpooled standard error from both groups’ standard deviations and sample sizes.
  • Compute the t statistic from summary statistics and keep the group order consistent.
  • Interpret the sign and size of the statistic in the context of the measured variable.
  • Recognize why a t statistic alone does not determine a test conclusion.

What the Two-Sample t Statistic Measures

In “Defining Both Population Means in Context,” you identified the two population means being compared. In this tutorial, we use the two samples’ summary statistics to calculate the test statistic for a test of whether those population means are equal. The statistic compares the observed difference in sample means with the difference proposed by the null hypothesis.

For the usual two-sample t test, the null hypothesis proposes no difference: \(H_0:\mu_1-\mu_2=0\). The sample difference, \(\bar{x}_1-\bar{x}_2\), will rarely be exactly zero. The t statistic tells us how many estimated standard errors that observed difference is above or below the null value.

Definition: The two-sample t test statistic is the observed difference in sample means minus the hypothesized difference, divided by the estimated standard error of the difference. For a null hypothesis of no difference, the hypothesized difference is zero.
$$ t=\frac{\bar{x}_1-\bar{x}_2-0}{SE_{\bar{x}_1-\bar{x}_2}} =\frac{\bar{x}_1-\bar{x}_2}{\sqrt{\frac{s_1^2}{n_1}+\frac{s_2^2}{n_2}}} $$

Here, \(\bar{x}_1\) and \(\bar{x}_2\) are the sample means, \(s_1\) and \(s_2\) are the sample standard deviations, and \(n_1\) and \(n_2\) are the sample sizes. As covered in “Standard Error for a Difference in Means,” calculate the two variance contributions separately and add them before taking the square root. This is the unpooled standard error.

The subtraction order matters. If group 1 is Treatment A and group 2 is Treatment B, the numerator is \(\bar{x}_A-\bar{x}_B\). A positive t statistic means the observed sample mean for A is greater than the observed sample mean for B. A negative t statistic means it is smaller. Reversing the group order reverses the sign of the statistic, but not its magnitude.

Calculate the Statistic from Summary Statistics

A reliable calculation has three parts: find the observed difference in means, find the standard error, and divide the first quantity by the second. Keep enough digits during the calculation and round the final t statistic only at the end.

1
Find the observed difference.
Calculate \(\bar{x}_1-\bar{x}_2\) in the order used in the hypotheses.
2
Calculate the standard error.
Square each sample standard deviation, divide by its own sample size, add those contributions, and take the square root.
3
Divide to get t.
For a null difference of zero, divide the observed difference by the standard error. Include the sign and interpret it using the group order.

The numerator and standard error have the same units, such as minutes or points. Those units cancel in the division, so \(t\) has no units. For example, \(t=2\) means the observed difference is two estimated standard errors above the null difference; it does not mean the groups differ by two measurement units.

Worked Examples

Worked Example: Comparing Two Focus Activities

Suppose a hypothetical study randomly assigns eligible adults to use either a guided focus activity or a standard quiet break before a timed task. The response is task-completion time in minutes. Let group 1 be the guided activity and group 2 the quiet break. The summary statistics are:

Group\(n\)\(\bar{x}\) (minutes)\(s\) (minutes)
Guided activity1642.08.0
Quiet break1447.07.5

State: Let \(\mu_1\) be the true mean task-completion time for eligible adults using the guided activity, and let \(\mu_2\) be the true mean for eligible adults taking the quiet break. To test \(H_0:\mu_1-\mu_2=0\), calculate the two-sample t statistic.

Plan: The response is quantitative, and each adult is in only one treatment group, so the groups are not paired. The adults were randomly selected from a much larger eligible population and randomly assigned to treatments; the sample is less than 10% of that population, and random assignment supports treating the groups as independent. Because both sample sizes are below 30, the task-time distributions should have no strong skewness or outliers. Suppose inspection of the study’s plots finds no such features.

Do: The observed difference, guided activity minus quiet break, is \(42.0-47.0=-5.0\) minutes. The standard error is:

$$ SE=\sqrt{\frac{8.0^2}{16}+\frac{7.5^2}{14}} =\sqrt{\frac{64}{16}+\frac{56.25}{14}} =\sqrt{4+4.0179} \approx 2.8316\text{ minutes} $$

Therefore,

$$ t=\frac{42.0-47.0-0}{2.8316} =\frac{-5.0}{2.8316} \approx -1.766 $$

Conclude: The observed difference in mean task-completion times is about 1.766 estimated standard errors below the null difference of zero, with guided activity minus quiet break as the subtraction order. This statistic alone does not determine whether there is convincing evidence of a difference; that judgment also requires the test’s reference distribution and p-value.

Worked Example: Comparing Two Water Filters

In a hypothetical randomized experiment, water samples are processed using Filter A or Filter B. The response is the reduction in turbidity, measured in a consistent unit. Use Filter A as group 1 and Filter B as group 2. The summary statistics are:

Group\(n\)\(\bar{x}\)\(s\)
Filter A2518.44.0
Filter B2016.13.5

The sample difference is \(18.4-16.1=2.3\) units of turbidity reduction. For a test of \(H_0:\mu_A-\mu_B=0\), calculate:

$$ SE=\sqrt{\frac{4.0^2}{25}+\frac{3.5^2}{20}} =\sqrt{\frac{16}{25}+\frac{12.25}{20}} =\sqrt{0.6400+0.6125} =\sqrt{1.2525} \approx 1.1192 $$
$$ t=\frac{18.4-16.1-0}{1.1192} =\frac{2.3}{1.1192} \approx 2.055 $$

The positive statistic indicates that Filter A’s sample mean reduction is above Filter B’s by about 2.055 estimated standard errors. The statistic is not a difference of 2.055 turbidity units: the sample means differ by 2.3 units, while \(t\) expresses that difference in standard-error units.

Worked Example: Comparing Two Package Designs

A hypothetical company randomly assigns independently selected shoppers to view one of two package designs. The measured response is the amount shoppers say they would pay, in dollars. Design X is group 1 and Design Y is group 2. Suppose the summary statistics are \(n_1=12\), \(\bar{x}_1=\$31.60\), \(s_1=\$5.40\), \(n_2=18\), \(\bar{x}_2=\$28.90\), and \(s_2=\$4.20\). The samples are each less than 10% of their respective target populations, the groups contain different shoppers, and plots show no strong skewness or outliers.

The observed difference in sample means is \(31.60-28.90=\$2.70\). Using \(H_0:\mu_X-\mu_Y=0\), the standard error is:

$$ SE=\sqrt{\frac{5.40^2}{12}+\frac{4.20^2}{18}} =\sqrt{\frac{29.16}{12}+\frac{17.64}{18}} =\sqrt{2.43+0.98} =\sqrt{3.41} \approx \$1.8466 $$

Then the test statistic is:

$$ t=\frac{31.60-28.90-0}{1.8466} =\frac{2.70}{1.8466} \approx 1.462 $$

The statistic is positive because Design X has the larger sample mean. The observed difference is about 1.462 estimated standard errors above zero. If the subtraction order were changed to Design Y minus Design X, the statistic would be approximately \(-1.462\); the evidence assessment would not change just because the labels were reversed.

What the Statistic Does—and Does Not—Tell You

A t statistic summarizes the sample difference relative to the variability expected in the difference in sample means. Holding the standard error fixed, a larger absolute difference produces a larger absolute t statistic. Holding the observed difference fixed, a larger standard error produces a smaller absolute t statistic.

The sign gives the direction of the observed difference according to the stated order. The absolute value gives its distance from the null value in estimated standard errors. A statistic near zero means the observed difference is close to the null value relative to its standard error. A statistic far from zero means it is farther from the null value in those units.

However, the statistic is not itself a p-value and does not, by itself, justify a conclusion about convincing evidence. To assess evidence, the t statistic must be compared with an appropriate t distribution, which depends on the degrees of freedom and the alternative hypothesis. The calculation here is the statistic component of the test.

Common Mistakes and AP Exam Tips

  • Putting the groups in the wrong order: If the hypotheses use \(\mu_1-\mu_2\), calculate \(\bar{x}_1-\bar{x}_2\), not the reverse. A reversed order changes the sign and can make the written interpretation inconsistent.
  • Forgetting to subtract the null value: The general numerator is observed difference minus hypothesized difference. For \(H_0:\mu_1-\mu_2=0\), this becomes \(\bar{x}_1-\bar{x}_2-0\).
  • Using the wrong standard error: Use \(\sqrt{s_1^2/n_1+s_2^2/n_2}\). Do not add the standard deviations, average them, or combine them into a pooled standard deviation for the standard unpooled two-sample t test.
  • Dividing by the sample standard deviation: The denominator is the standard error of the difference in means, not \(s_1\), \(s_2\), or a standard deviation from just one group.
  • Dropping the sign: Report a negative statistic when the first sample mean is smaller than the second. The sign communicates direction; do not replace \(t\) with \(|t|\) unless a task specifically asks for the absolute value.
  • Giving t measurement units: The mean difference and standard error use the same units, so the units cancel. Report the means and standard error with their measurement units, but the t statistic without units.
  • Calling the statistic a conclusion: “\(t=2.055\), so there is convincing evidence” is incomplete. A full-credit response identifies what the statistic measures and uses the appropriate p-value and context before making an evidence conclusion.

For a clear AP response, show the standard-error formula with each group’s values substituted, show the numerator in the stated order, and report the resulting statistic with an appropriate interpretation. Keep the t statistic distinct from the p-value and from the original difference in measurement units.

Key takeaway: For a two-sample t test of \(H_0:\mu_1-\mu_2=0\), calculate \(t=(\bar{x}_1-\bar{x}_2)/\sqrt{s_1^2/n_1+s_2^2/n_2}\). The sign follows the group order, and the statistic measures the observed difference in estimated standard-error units.

Check Your Understanding

Use the two-sample t statistic formula and keep the group order clear in each response.

  1. Group 1 has \(n_1=16\), \(\bar{x}_1=24\), and \(s_1=4\). Group 2 has \(n_2=25\), \(\bar{x}_2=21\), and \(s_2=5\). Calculate the standard error and t statistic for \(H_0:\mu_1-\mu_2=0\).
  2. If \(\bar{x}_1-\bar{x}_2\) is negative and the null difference is zero, what is the sign of the t statistic? Explain what the sign says about the sample means.
  3. Why must the standard error use both \(s_1^2/n_1\) and \(s_2^2/n_2\)?
  4. A student reports \(t=1.8\) and says the groups’ sample means differ by 1.8 minutes. What is wrong with that interpretation?
  5. Suppose the group labels are swapped but all sample statistics stay the same. What happens to the sign and absolute value of the t statistic?