Tutorials › AP Statistics › Sampling Distribution of x-bar1 Minus x-bar2

Two-sample t confidence intervals · Tutorial 703 of 1000

Sampling Distribution of x-bar1 Minus x-bar2

Describe the center, spread, and shape of the sampling distribution of the difference between two independent sample means.

Intermediate 10 min read

What You'll Learn

  • Define the random variable \(\bar{x}_1-\bar{x}_2\) for two independent samples
  • Find its expected value using the difference between population means
  • Calculate its standard deviation from the population standard deviations and sample sizes
  • Decide when the sampling distribution is normal or approximately normal
  • Check independence, the 10% condition, and sample-size conditions in context
  • Explain how the sample sizes affect the variability of the difference

From Two Groups to a Sampling Distribution

In “Comparing Two Means: The Parameter of Interest,” the population quantity \(\mu_1-\mu_2\) was defined as population 1’s mean minus population 2’s mean. In “Identifying Two Independent Samples,” we learned to use the study design to decide whether a two-sample comparison is appropriate. Now consider what would happen if we repeatedly took new independent samples from the two populations and calculated the difference between their sample means each time.

That difference would vary from sample to sample. The probability distribution of all possible values of \(\bar{x}_1-\bar{x}_2\), under a specified sampling process, is its sampling distribution. Describing its center, spread, and shape helps us understand how close a sample difference is likely to be to the population difference.

Definition: Let \(\bar{x}_1\) and \(\bar{x}_2\) be the sample means from two independent samples. The random variable \(D=\bar{x}_1-\bar{x}_2\) is the difference between the sample means, with sample 1’s mean first and sample 2’s mean second.

The order matters. If population 1’s mean is greater than population 2’s mean, the center of the sampling distribution is positive. Reversing the sample order changes the sign of the difference, but not its spread.

Center and Spread

For independent samples, the center of the sampling distribution of \(\bar{x}_1-\bar{x}_2\) is the difference between the population means, \(\mu_1-\mu_2\). The sample means are centered at their respective population means, so subtracting one sample mean from the other centers the difference at the corresponding difference in population means.

The spread depends on both populations and both sample sizes. Let \(\sigma_1\) and \(\sigma_2\) be the population standard deviations. The variance of the difference is the sum of the variances of the two sample means when the samples are independent. Taking the square root gives the standard deviation of the sampling distribution.

$$ \text{Mean of }D=\mu_1-\mu_2 \qquad \text{SD of }D=\sqrt{\frac{\sigma_1^2}{n_1}+\frac{\sigma_2^2}{n_2}} $$

Here \(n_1\) and \(n_2\) are the sample sizes. This standard deviation describes the typical distance between a sample difference \(\bar{x}_1-\bar{x}_2\) and the population difference \(\mu_1-\mu_2\), across repeated sampling under the same conditions. It is measured in the same units as the original quantitative variable.

Notice that the variances add; we do not subtract the standard deviations. Each sample mean varies around its own population mean. When the two samples are independent, those two sources of variation contribute separately to the variation in their difference.

Key distinction: The formula above describes the theoretical standard deviation of the sampling distribution and uses the population standard deviations \(\sigma_1\) and \(\sigma_2\). In practice, those population values are usually unknown. The estimated standard deviation used in a two-sample t interval is called the standard error and is the subject of the next tutorial.

When Is the Sampling Distribution Approximately Normal?

If both population distributions are normal, then each sample mean has a normal distribution. For independent samples, their difference is also normally distributed. So, with normally distributed populations, the sampling distribution of \(\bar{x}_1-\bar{x}_2\) is normal even when the sample sizes are small.

If a population is not normal, the Central Limit Theorem explains why its sample mean can still be approximately normal when the sample size is sufficiently large. For the difference of two sample means to be approximately normal by this reasoning, each sample mean should be approximately normal. Check the sample sizes separately: a large combined total does not make up for one very small sample.

A common introductory guideline is that a sample size of at least 30 may be sufficient for a sample mean when the population is not normal. This is a rule of thumb, not a guarantee. Strong skewness or pronounced outliers may require larger samples. With small samples, look for populations that are approximately normal, without strong skewness or outliers.

Conditions for the usual sampling-distribution description:
  • Independent samples: The groups have no pairing or other design-based link, as described in “Identifying Two Independent Samples.” The sampling or assignment process must support treating the two sample means as independent.
  • Randomness and the 10% condition: For sampling without replacement, each sample should be no more than 10% of its population when independence within the sample is approximated. Check this separately for each population.
  • Normal shape: Both population distributions are normal, or both sample sizes are sufficiently large for their sample means to be approximately normal. Consider skewness and outliers when sample sizes are small.

These conditions support a model for repeated sampling; they do not make the observed sample difference equal to the population difference. A sample difference can be above or below the population difference because of sampling variability.

Worked Examples: Center, Spread, and Shape

Worked Example: Comparing Two Normal Populations

Imagine two independent random samples of seedlings grown under different light conditions. Let population 1 be seedlings under blue light and population 2 be seedlings under white light. Suppose the plant heights in the two populations are normally distributed, with \(\mu_1=72\) centimeters and \(\sigma_1=10\) centimeters for blue light, and \(\mu_2=68\) centimeters and \(\sigma_2=8\) centimeters for white light. The sample sizes are \(n_1=25\) and \(n_2=16\). Assume each sample is no more than 10% of its population.

Let \(D=\bar{x}_1-\bar{x}_2\), the blue-light sample mean height minus the white-light sample mean height. The samples are independent, and the populations are normal, so the sampling distribution of \(D\) is normal. The 10% comparison supports treating observations within each sample as independent.

The center is:

$$ \mu_1-\mu_2=72-68=4\text{ centimeters} $$

The standard deviation is:

$$ \sqrt{\frac{10^2}{25}+\frac{8^2}{16}} =\sqrt{\frac{100}{25}+\frac{64}{16}} =\sqrt{4+4} =\sqrt{8} \approx 2.8284\text{ centimeters} $$

The variance calculation checks the result: \(4+4=8\), and the square root of 8 is approximately 2.8284. Thus \(D\) is normally distributed with mean 4 centimeters and standard deviation about 2.8284 centimeters. The positive center reflects that the blue-light population mean exceeds the white-light population mean by 4 centimeters; individual sample differences will vary around that center.

Worked Example: Large Samples From Skewed Populations

A community garden compares the weekly harvest mass, in kilograms, from two kinds of planting beds. Suppose population 1 has mean \(\mu_1=32\) kilograms and standard deviation \(\sigma_1=12\) kilograms; population 2 has mean \(\mu_2=28\) kilograms and standard deviation \(\sigma_2=15\) kilograms. The harvest distributions are strongly right-skewed. Independent random samples of \(n_1=40\) and \(n_2=50\) beds are selected. Each sample is less than 10% of its respective population.

Let \(D=\bar{x}_1-\bar{x}_2\), the sample mean harvest for population 1 minus that for population 2. The sampling is random, the groups are independent, and the 10% condition is satisfied for both populations. Both sample sizes meet a common guideline for using the Central Limit Theorem, but strong right-skewness means this does not guarantee that the sample means are approximately normal; larger samples or more information about the distributions may be needed. Thus, the information given does not establish that the difference is approximately normal.

Its center is \(32-28=4\) kilograms. Its standard deviation is:

$$ \sqrt{\frac{12^2}{40}+\frac{15^2}{50}} =\sqrt{\frac{144}{40}+\frac{225}{50}} =\sqrt{3.6+4.5} =\sqrt{8.1} \approx 2.8460\text{ kilograms} $$

As a check, the two variance contributions are 3.6 and 4.5, which sum to 8.1; the square root is approximately 2.8460. Thus the difference in sample mean harvests is centered at 4 kilograms, with a standard deviation of about 2.8460 kilograms, but the information given does not establish whether its sampling distribution is approximately normal. The conclusion about shape uses both sample sizes, not just their total of 90.

Worked Example: A Small Sample With a Strongly Skewed Population

Suppose independent random samples compare the time, in hours, that two types of rechargeable lanterns operate before needing a new battery. For type 1, assume \(\mu_1=5.2\) hours and \(\sigma_1=0.8\) hours, with \(n_1=8\). For type 2, assume \(\mu_2=4.7\) hours and \(\sigma_2=1.6\) hours, with \(n_2=9\). The type 2 operating times have a strong right skew. Assume both samples are less than 10% of their respective populations.

The samples are independent and the 10% condition is met. The center of \(D=\bar{x}_1-\bar{x}_2\) is \(5.2-4.7=0.5\) hours. Its standard deviation is:

$$ \sqrt{\frac{0.8^2}{8}+\frac{1.6^2}{9}} =\sqrt{\frac{0.64}{8}+\frac{2.56}{9}} =\sqrt{0.08+0.2844\ldots} \approx 0.6037\text{ hours} $$

The variance contributions add to approximately 0.3644, and the square root of 0.3644 is approximately 0.6037. The standard deviation calculation is valid, but the shape of the sampling distribution is not established as approximately normal: both sample sizes are small, and one population is strongly skewed. The information given does not justify a normal model for \(D\). Independence alone does not guarantee a normal sampling distribution.

How Sample Size Affects the Difference

The formula shows why increasing either sample size reduces the spread of the sampling distribution. As \(n_1\) increases, \(\sigma_1^2/n_1\) decreases; as \(n_2\) increases, \(\sigma_2^2/n_2\) decreases. If the samples remain independent and the population standard deviations stay fixed, taking larger samples makes the sample difference less variable from one repetition to another.

The sample sizes need not be equal. Each population contributes its own variance term, so the effect of increasing \(n_1\) depends on \(\sigma_1\), and the effect of increasing \(n_2\) depends on \(\sigma_2\). The formula accounts for both contributions rather than treating the total sample size as a single number.

The center does not change just because the sample sizes change: it remains \(\mu_1-\mu_2\). Larger samples reduce spread around that same center. They do not, by themselves, change the population means or guarantee that a particular observed difference equals the center.

Common Mistakes and AP Exam Tips

  • Using the wrong order: If the random variable is \(\bar{x}_1-\bar{x}_2\), its center is \(\mu_1-\mu_2\), in that same order. State what group 1 and group 2 represent.
  • Subtracting variances: For independent sample means, the variances add: \(\sigma_1^2/n_1+\sigma_2^2/n_2\). Then take the square root to get the standard deviation.
  • Using standard deviations without squaring them: The terms in the formula are \(\sigma_1^2/n_1\) and \(\sigma_2^2/n_2\), not \(\sigma_1/n_1\) and \(\sigma_2/n_2\).
  • Checking only the combined sample size: For an approximate-normal argument based on large samples, assess \(n_1\) and \(n_2\) separately. A very large sample in one group does not automatically fix a small sample from a strongly skewed population in the other.
  • Assuming random samples guarantee normality: Randomness supports the sampling process and independence checks; it does not make a skewed population normal. For small samples, the population shape matters.
  • Confusing theoretical standard deviation with standard error: The sampling-distribution formula uses population standard deviations. When these are unknown and estimated from data, the corresponding estimated spread is called the standard error.

For full-credit communication, define the difference in the stated order, report the center and spread with units, and explain why the shape is normal, approximately normal, or not established. Name the design facts that support independence and address the 10% condition when sampling without replacement. Avoid saying the sample difference “is” the population difference; the sampling distribution describes how sample results vary around it.

Key takeaway: For independent samples, the sampling distribution of \(\bar{x}_1-\bar{x}_2\) is centered at \(\mu_1-\mu_2\), with standard deviation \(\sqrt{\sigma_1^2/n_1+\sigma_2^2/n_2}\). Its shape is normal for normal populations and approximately normal when both sample means are approximately normal.

Check Your Understanding

Use the sampling-distribution ideas in each question. Include units and explain any condition relevant to your answer.

  1. Independent samples have \(n_1=36\), \(n_2=49\), \(\sigma_1=6\), and \(\sigma_2=7\). What is the standard deviation of \(\bar{x}_1-\bar{x}_2\)?
  2. If \(\mu_1=18\) and \(\mu_2=23\), what is the center of the sampling distribution of \(\bar{x}_1-\bar{x}_2\), and what does its sign mean?
  3. Two independent random samples are drawn without replacement from populations of 300 and 500 individuals. The sample sizes are 20 and 45. Does each sample meet the 10% condition?
  4. Both populations are normal, and the samples are independent. What can you say about the shape of the sampling distribution, even if the sample sizes are small?
  5. One sample size is 12 and its population is strongly right-skewed; the other sample size is 80 and its population is approximately normal. Is an approximately normal shape for the difference justified by the information given? Explain.