Tutorials › AP Statistics › Pooled Two-Sample t Test and Why AP Avoids It

Two-sample t hypothesis tests · Tutorial 737 of 1000

Pooled Two-Sample t Test and Why AP Avoids It

See how pooling works, what its equal-variance assumption means, and why the unpooled two-sample t test is the safer AP Statistics default.

Intermediate 9 min read

What You'll Learn

  • Explain the equal-variance assumption behind a pooled two-sample t test.
  • Calculate the pooled variance, pooled standard deviation, standard error, and degrees of freedom.
  • Distinguish a null hypothesis of equal means from an assumption of equal variances.
  • Describe why AP Statistics defaults to the unpooled Welch procedure.
  • Recognize when pooled and unpooled calculations can differ and what that means for a conclusion.

Pooling Makes an Extra Assumption

In “Comparing Two Means Using Boxplots First,” you used the two groups’ distributions as an initial check before inference. The usual AP two-sample t test then compares the means of independent groups using each group’s sample variance separately. A pooled two-sample t test takes a different approach: it combines the sample variances into one estimate of a shared population variance.

That combination is useful only if a particular assumption is reasonable: the two populations have the same variance. This equal-variance assumption is about the populations’ variability, not about their means. A test can ask whether population means are equal while separately assuming that population variances are equal.

Definition: A pooled two-sample t test assumes the two independent populations have a common variance, so one shared variance can be estimated from both samples. The test still concerns the difference in population means, \(\mu_1-\mu_2\).

Pooling gives more weight to the sample variance from the group with more degrees of freedom. If the populations really have a common variance, using information from both groups can estimate that variance efficiently. If their variances differ, however, the combined estimate may not represent either group well. The resulting standard error—and therefore the test statistic and p-value—can be misleading, especially when the sample sizes are unequal.

How the Pooled Calculation Works

For independent samples of sizes \(n_1\) and \(n_2\), with sample standard deviations \(s_1\) and \(s_2\), the pooled variance is a weighted average of the two sample variances. The weights are the groups’ degrees of freedom, \(n_1-1\) and \(n_2-1\).

$$ s_p^2=\frac{(n_1-1)s_1^2+(n_2-1)s_2^2}{n_1+n_2-2} $$

The pooled standard deviation is \(s_p=\sqrt{s_p^2}\). Under the equal-variance assumption, the standard error for the difference in sample means and the degrees of freedom are:

$$ SE_{\text{pooled}}=s_p\sqrt{\frac{1}{n_1}+\frac{1}{n_2}}, \qquad df=n_1+n_2-2 $$

For a null hypothesis of no difference, the pooled test statistic is \(t=(\bar{x}_1-\bar{x}_2)/SE_{\text{pooled}}\). The group order determines the sign, just as in the unpooled test. The formula differs because pooling replaces the two separate variance contributions with one shared estimate.

Important distinction: \(H_0:\mu_1-\mu_2=0\) states that the population means are equal. It does not state that the population variances are equal. The pooled test adds equal population variances as a separate assumption.

Worked Examples

Worked Example: Combining Two Variance Estimates

A fictional engineering class compares the breaking strength, in newtons, of two independently sampled types of model bridge joint. The Type A sample has \(n_1=8\) and \(s_1=3\) newtons; the Type B sample has \(n_2=20\) and \(s_2=6\) newtons. Calculate the pooled variance and pooled standard deviation, then compare the pooled and unpooled standard errors.

Calculate the pooled variance: The degrees of freedom are 7 for Type A and 19 for Type B. Thus,

$$ s_p^2 =\frac{7(3^2)+19(6^2)}{8+20-2} =\frac{63+684}{26} =\frac{747}{26} \approx 28.7308\text{ newtons}^2. $$

Taking the square root gives \(s_p=\sqrt{28.7308}\approx5.3601\) newtons. The pooled standard error is

$$ SE_{\text{pooled}} =5.3601\sqrt{\frac18+\frac1{20}} =5.3601\sqrt{0.175} \approx2.2423\text{ newtons}. $$

Compare with the unpooled standard error: The unpooled calculation keeps the variance contributions separate:

$$ SE_{\text{unpooled}} =\sqrt{\frac{3^2}{8}+\frac{6^2}{20}} =\sqrt{1.125+1.8} =\sqrt{2.925} \approx1.7103\text{ newtons}. $$

The pooled standard error is larger in this example. The smaller sample has the smaller sample standard deviation, while the larger sample has the larger one; the pooled calculation combines them and applies the common-variance model. The unpooled calculation reflects each group’s observed variance and sample size directly. Which standard error is justified depends on the assumptions, not on choosing whichever gives a preferred result.

This calculation alone does not establish that the populations have equal variances. That claim concerns population parameters; two sample standard deviations are only estimates and can differ by chance.

Worked Example: A Pooled Test and Its Assumption

A fictional materials lab randomly samples two independent types of coating and measures the drying time, in minutes. Suppose the summaries are \(n_1=12\), \(\bar{x}_1=31.1027\), \(s_1=4\) for Type 1 and \(n_2=15\), \(\bar{x}_2=27.0000\), \(s_2=5\) for Type 2. For illustration, carry out a pooled two-sided test at \(\alpha=0.05\), assuming a common population variance. The sample means are reported to enough decimal places to make the calculation below approximately \(t=4/\sqrt{3}\).

1
State.
Let \(\mu_1\) and \(\mu_2\) be the true mean drying times for the two coating populations. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\). The equal-variance assumption is that the two populations have a common variance.
2
Plan and check conditions.
The setting specifies random samples and independent groups, so the random-design and independent-samples conditions are met. Assume each sample is less than 10% of its population, satisfying the 10% condition for sampling without replacement. With sample sizes 12 and 15, inspect each group’s distribution carefully; suppose the lab’s plots show no strong skewness or outliers. Finally, the pooled method requires equal population variances. Similar sample standard deviations, 4 and 5 minutes, may make that assumption seem plausible, but they cannot prove it.
3
Do.
First calculate the pooled variance and standard deviation:
$$ s_p^2 =\frac{11(4^2)+14(5^2)}{25} =\frac{176+350}{25} =\frac{526}{25} =21.04\text{ minutes}^2, \qquad s_p=\sqrt{21.04}\approx4.5869\text{ minutes}. $$

The pooled standard error is

$$ SE_{\text{pooled}} =\sqrt{21.04\left(\frac1{12}+\frac1{15}\right)} =\sqrt{21.04(0.15)} =\sqrt{3.156} \approx1.7765\text{ minutes}. $$

The degrees of freedom are \(12+15-2=25\). Using the stated sample means, the test statistic is approximately

$$ t=\frac{31.1027-27.0000}{1.7765}\approx2.3094. $$

For a two-sided test, the p-value is the probability, assuming \(H_0\) is true, of getting a t statistic at least as far from zero as \(2.3094\), using a t distribution with 25 degrees of freedom. A calculator gives \(p\approx0.02947\), rounded.

4
Conclude in context.
Because \(0.02947<0.05\), reject \(H_0\). If the equal-variance assumption and the other conditions are reasonable, these data provide convincing evidence that the true mean drying times differ between the two coating populations. This conclusion is conditional on the pooled method’s extra assumption.

For AP Statistics, the usual next step would be to use the unpooled Welch two-sample t test instead. Here its standard error is \(\sqrt{4^2/12+5^2/15}=\sqrt{3}\approx1.7321\) minutes, rather than the pooled \(1.7765\) minutes. The procedures can give different test statistics and p-values; AP’s standard conclusion should come from the unpooled procedure, not from selecting whichever result is more favorable.

Worked Example: Equal Sample Sizes Can Hide a Difference

A fictional school compares the time, in minutes, that students in two independent groups take to complete a short practice activity. Each group has 10 students. The sample standard deviations are \(s_1=4\) and \(s_2=5\) minutes. Compare the pooled and unpooled standard errors.

The pooled variance is

$$ s_p^2 =\frac{9(4^2)+9(5^2)}{18} =\frac{144+225}{18} =20.5\text{ minutes}^2. $$

Therefore, the pooled standard error is

$$ SE_{\text{pooled}} =\sqrt{20.5\left(\frac1{10}+\frac1{10}\right)} =\sqrt{4.1} \approx2.0248\text{ minutes}. $$

The unpooled standard error is

$$ SE_{\text{unpooled}} =\sqrt{\frac{16}{10}+\frac{25}{10}} =\sqrt{4.1} \approx2.0248\text{ minutes}. $$

The standard errors match because the sample sizes are equal: with equal denominators, pooling the two sample variances and then multiplying by \(1/n_1+1/n_2\) gives the same variance estimate as adding their separate contributions. This does not make the equal-variance assumption true. The procedures can still use different degrees of freedom, and with unequal sample sizes their standard errors generally need not match.

Why AP Uses the Unpooled Test by Default

The unpooled two-sample t test, also called Welch’s test, estimates each group’s variance contribution separately:

$$ SE_{\text{unpooled}} =\sqrt{\frac{s_1^2}{n_1}+\frac{s_2^2}{n_2}}. $$

As covered in “The Two-Sample t Test Statistic” and “Computing Two-Sample Degrees of Freedom With the Welch Formula,” the unpooled test uses Welch degrees of freedom. It does not require the two population variances to be equal. This makes it a reliable default when the population variances are unknown—which is the usual situation in an introductory statistics problem.

The pooled test can perform well when the equal-variance assumption is appropriate. But a student usually cannot verify that assumption from a pair of sample standard deviations or boxplots. A visual comparison can flag a striking difference in spread, as in “Comparing Two Means Using Boxplots First,” but a lack of an obvious difference is not proof that the population variances are equal. AP avoids making pooled inference the default because the unpooled method does not need that extra assumption and remains appropriate across a wider range of settings.

Key takeaway: Pooling combines sample variances under the assumption of equal population variances. AP Statistics uses the unpooled Welch two-sample t test by default because it keeps the groups’ variance contributions separate and does not require that assumption.

Common Mistakes and AP Exam Tips

  • Confusing equal means with equal variances: The null hypothesis concerns \(\mu_1-\mu_2\). Equal population variances are a separate assumption required by the pooled method.
  • Pooling standard deviations directly: The pooled quantity is a weighted average of the sample variances, \(s_1^2\) and \(s_2^2\), not a simple average of \(s_1\) and \(s_2\). Take the square root only after calculating the pooled variance.
  • Using pooled degrees of freedom for Welch’s test: The pooled test has \(n_1+n_2-2\) degrees of freedom. The unpooled test uses Welch degrees of freedom (or a specified conservative alternative), as explained earlier in this course.
  • Claiming sample spreads prove population spreads are equal: Similar sample standard deviations do not establish equal population variances. State that pooling assumes equal variances; do not claim the assumption has been proven by the sample.
  • Calling the pooled method the AP default: For the standard AP two-sample t test, use the unpooled method. On a TI-84, choose “Pooled: No” for 2-SampTTest unless a problem explicitly directs you to use a pooled calculation.
  • Giving a conclusion without its assumption: If asked to interpret a pooled test, connect the inference to the equal-variance assumption. For the standard AP response, use the unpooled test and conclude in context using its p-value and the stated significance level.

A full-credit response identifies the test and hypotheses, checks the study design and distribution conditions, and uses the correct standard error and degrees of freedom for the procedure. If pooling is specifically requested, name the equal-variance assumption. If it is not requested, the AP default is the unpooled Welch test.

Check Your Understanding

Answer each question using the ideas and calculations in this tutorial.

  1. What population assumption allows a pooled two-sample t test to combine the sample variances?
  2. Does \(H_0:\mu_1-\mu_2=0\) assert that the two population variances are equal? Explain.
  3. For the bridge-joint example, why does the pooled standard error differ from the unpooled standard error?
  4. In the coating example, what does the pooled test’s p-value measure, and what assumption is its conclusion conditional on?
  5. Why is the unpooled Welch test the standard AP Statistics choice when a problem does not specify pooling?