Tutorials › AP Statistics › Sampling Distribution of the Difference in Sample Proportions

Two-proportion confidence intervals · Tutorial 503 of 1000

Sampling Distribution of the Difference in Sample Proportions

Describe the center, spread, and shape of the sampling distribution of \(\hat{p}_1-\hat{p}_2\), and connect its mean to \(p_1-p_2\).

Intermediate 9 min read

What You'll Learn

  • Identify the mean of the sampling distribution of \(\hat{p}_1-\hat{p}_2\).
  • Calculate its standard deviation from the two population proportions and sample sizes.
  • Check the conditions needed for an approximately Normal sampling distribution.
  • Distinguish the distribution’s theoretical standard deviation from a value calculated using sample data.
  • Explain what the distribution says about repeated differences between sample proportions.

From One Estimate to a Sampling Distribution

In Point Estimate for the Difference of Two Proportions, you learned that \(\hat{p}_1-\hat{p}_2\) estimates the population difference \(p_1-p_2\). But a new pair of samples would usually produce a different estimate. The sampling distribution describes how the statistic \(\hat{p}_1-\hat{p}_2\) varies across all possible pairs of samples of the same sizes, selected in the same way.

The group order stays fixed: Group 1 minus Group 2. Each sample proportion varies from sample to sample, so their difference varies too. The sampling distribution has a center, a spread, and a shape. Understanding these features helps describe how much the estimate typically varies and when a Normal model is reasonable.

Definition: The sampling distribution of \(\hat{p}_1-\hat{p}_2\) is the distribution of the differences in sample proportions from all possible pairs of samples of sizes \(n_1\) and \(n_2\), selected under the same conditions.

Center and Spread

The mean of the sampling distribution is the true population difference, \(p_1-p_2\). In repeated sampling, \(\hat{p}_1-\hat{p}_2\) is centered at the difference it estimates. This is why the statistic is called an unbiased estimator of \(p_1-p_2\): over many repetitions, its average value equals the population difference.

The spread of the sampling distribution is measured by the standard deviation of \(\hat{p}_1-\hat{p}_2\). For independent samples, the formula combines the variability of each sample proportion. The standard deviation is in proportion units; for example, \(0.05\) corresponds to 5 percentage points.

$$ \mu_{\hat{p}_1-\hat{p}_2}=p_1-p_2 \qquad\text{and}\qquad \sigma_{\hat{p}_1-\hat{p}_2} = \sqrt{\frac{p_1(1-p_1)}{n_1}+\frac{p_2(1-p_2)}{n_2}} $$

The spread depends on both population proportions and both sample sizes. A larger sample size generally makes its sample proportion less variable. The true proportions also matter: a sample proportion tends to be most variable when its population proportion is near \(0.50\), and less variable when it is near 0 or 1.

This formula describes the theoretical standard deviation when the population proportions are specified. In real inference problems, those proportions are usually unknown. A later tutorial addresses how to estimate the spread from sample data. For now, focus on what the sampling distribution would be if \(p_1\) and \(p_2\) were known.

Key idea: The mean of the sampling distribution is \(p_1-p_2\). For independent samples, its standard deviation is the square root of the sum of the two sample-proportion variances. Keep the Group 1-minus-Group 2 order throughout.

When Is the Shape Approximately Normal?

The sampling distribution of a difference in sample proportions is not automatically Normal. A Normal approximation is reasonable when each group has enough expected successes and failures. These checks concern the population proportions and the sample sizes, not merely the counts observed in one particular pair of samples.

Conditions:
  • Random: Each sample should be a random sample from its population, or the data should come from an appropriate randomized process.
  • Independent groups: The two samples should be selected independently of one another.
  • 10% condition: If sampling without replacement from a finite population, each sample size should be no more than 10% of its population size. Check this separately for each group.
  • Large Counts: For each group, the expected number of successes and failures should each be at least 10: \(n_1p_1 \ge 10\), \(n_1(1-p_1) \ge 10\), \(n_2p_2 \ge 10\), and \(n_2(1-p_2) \ge 10\).

When these conditions are met, the sampling distribution of \(\hat{p}_1-\hat{p}_2\) is approximately Normal, centered at \(p_1-p_2\), with the standard deviation given above. If a Large Counts check fails, do not claim the distribution is approximately Normal based on this condition. It may be skewed or noticeably discrete instead.

Worked Examples

Worked Example: Center, Spread, and Shape for Two Health Surveys

Setting: Imagine two independent random samples of adults from two towns. Suppose the true proportion who meet a specified activity goal is \(p_1=0.60\) in Town 1 and \(p_2=0.52\) in Town 2. The sample sizes are \(n_1=200\) and \(n_2=250\). Assume each town has at least 10 times its sample size in adults.

Center: The mean difference is the Town 1 proportion minus the Town 2 proportion.

$$ \mu_{\hat{p}_1-\hat{p}_2} = p_1-p_2 = 0.60-0.52 = 0.08 $$

Spread: Substitute both population proportions and sample sizes into the standard-deviation formula.

$$ \sigma_{\hat{p}_1-\hat{p}_2} = \sqrt{\frac{0.60(0.40)}{200}+\frac{0.52(0.48)}{250}} = \sqrt{0.001200+0.0009984} = \sqrt{0.0021984} \approx 0.0469 $$

The arithmetic can be checked by calculating the two variance contributions separately: \(0.60(0.40)/200=0.001200\), and \(0.52(0.48)/250=0.0009984\). Their sum is \(0.0021984\), whose square root is approximately \(0.0469\). The standard deviation is about 4.69 percentage points.

Shape and conditions: The samples are stated to be random and independent. The 10% condition is met by the population-size assumption for each town. The expected counts are \(200(0.60)=120\) successes and \(200(0.40)=80\) failures for Town 1, and \(250(0.52)=130\) successes and \(250(0.48)=120\) failures for Town 2. All four counts are at least 10, so the sampling distribution is approximately Normal.

Thus the sampling distribution is approximately Normal with mean \(0.08\) and standard deviation \(0.0469\). Across repeated pairs of samples of these sizes, the difference in sample proportions is centered at 8 percentage points. Its typical distance from that center, measured by the standard deviation, is about 4.69 percentage points.

Worked Example: A Negative Center with Unequal Sample Sizes

Setting: In a fictional study of two independent customer populations, suppose \(p_1=0.30\) and \(p_2=0.45\) are the true proportions who would choose a particular delivery option. Random samples have sizes \(n_1=100\) and \(n_2=120\). Each population is at least 10 times its sample size.

Center: Keep the stated order, Group 1 minus Group 2.

$$ \mu_{\hat{p}_1-\hat{p}_2} = p_1-p_2 = 0.30-0.45 = -0.15 $$

The negative mean says that the true proportion in Group 1 is 15 percentage points lower than the true proportion in Group 2. The sign follows from the subtraction order, not from a change in how the sampling distribution is defined.

Spread:

$$ \sigma_{\hat{p}_1-\hat{p}_2} = \sqrt{\frac{0.30(0.70)}{100}+\frac{0.45(0.55)}{120}} = \sqrt{0.002100+0.0020625} = \sqrt{0.0041625} \approx 0.0645 $$

Checking the calculation by parts, the first contribution is \(0.21/100=0.002100\), and the second is \(0.2475/120=0.0020625\). Their sum is \(0.0041625\); its square root is approximately \(0.0645\), or 6.45 percentage points.

Shape and conditions: The samples are random and independent, and the stated population sizes satisfy the 10% condition. The expected success and failure counts are 30 and 70 in Group 1, and 54 and 66 in Group 2. Each is at least 10, so the sampling distribution is approximately Normal.

The distribution is therefore approximately Normal with mean \(-0.15\) and standard deviation \(0.0645\). It is centered below zero because Group 1’s population proportion is lower. Its spread describes variation around that negative center; it does not change the group order or turn the mean into a sample result.

Worked Example: A Large Counts Check That Fails

Setting: Suppose two independent random samples are taken from large populations. In Group 1, the true proportion with a specified characteristic is \(p_1=0.08\), and the sample size is \(n_1=50\). In Group 2, \(p_2=0.40\) and \(n_2=60\). Assume both samples are at most 10% of their populations.

Find the center:

$$ \mu_{\hat{p}_1-\hat{p}_2} = p_1-p_2 = 0.08-0.40 = -0.32 $$

Find the theoretical standard deviation:

$$ \sigma_{\hat{p}_1-\hat{p}_2} = \sqrt{\frac{0.08(0.92)}{50}+\frac{0.40(0.60)}{60}} = \sqrt{0.001472+0.004000} = \sqrt{0.005472} \approx 0.0740 $$

The components check as \(0.0736/50=0.001472\) and \(0.24/60=0.004000\). Adding them gives \(0.005472\), and the square root is about \(0.0740\).

Check the shape condition: The expected counts are 4 successes and 46 failures in Group 1, and 24 successes and 36 failures in Group 2. Group 1 has fewer than 10 expected successes, so the Large Counts condition fails. Even though the random, independence, and 10% conditions are satisfied, we should not use the usual Normal approximation based on these checks.

The mean and theoretical standard deviation are still as calculated, but those two values alone do not establish a Normal shape. Here the small expected success count in Group 1 warns that the sampling distribution may not be well approximated by a Normal curve.

Common Mistakes and AP Exam Tip

  • Using the sample difference as the mean: The mean of the sampling distribution is \(p_1-p_2\), not the observed \(\hat{p}_1-\hat{p}_2\). The sample difference is one possible value of the statistic.
  • Subtracting the population proportions in the wrong order: If the statistic is Group 1 minus Group 2, its mean is \(p_1-p_2\). Reversing the order changes the sign.
  • Leaving out one group’s variability: The standard deviation uses a contribution from each independent sample. Include both terms under the square root.
  • Checking only the total sample size: The Large Counts condition must be checked separately for successes and failures in each group. A large combined sample does not make up for a small expected count in one group.
  • Checking observed counts instead of expected counts: For this theoretical shape condition, check \(n_ip_i\) and \(n_i(1-p_i)\), not just the success and failure counts in one observed sample.
  • Calling the spread a guaranteed error: A standard deviation describes the distribution of repeated sample differences. It does not say that every sample difference will be within one standard deviation of the mean.
AP Exam Tip: State the center as \(p_1-p_2\), calculate both terms in the standard deviation, and show each group’s Large Counts checks. When the conditions are met, describe the shape as approximately Normal and name its center and standard deviation. If a condition fails, say so rather than asserting a Normal model.

Key Takeaway

The sampling distribution describes the values of \(\hat{p}_1-\hat{p}_2\) across repeated pairs of samples. Its mean is the population difference \(p_1-p_2\); for independent samples, its standard deviation is calculated from both population proportions and both sample sizes. Randomness, independence, the 10% condition, and Large Counts checks determine whether its shape is approximately Normal.

Key takeaway: Center the sampling distribution at \(p_1-p_2\), measure its spread with \(\sqrt{p_1(1-p_1)/n_1+p_2(1-p_2)/n_2}\), and check the conditions before using an approximately Normal shape.

Check Your Understanding

Use the group order given in each question. Show your calculations and explain what the sampling-distribution features mean.

  1. Suppose \(p_1=0.25\), \(p_2=0.35\), \(n_1=160\), and \(n_2=140\). Find the mean and standard deviation of \(\hat{p}_1-\hat{p}_2\).
  2. For the values in Question 1, calculate the expected success and failure counts in each group. Is the Large Counts condition satisfied?
  3. In a setting with \(p_1=0.70\), \(n_1=80\), \(p_2=0.20\), and \(n_2=100\), what is the mean of the sampling distribution? Interpret its sign using the Group 1-minus-Group 2 order.
  4. Why must the 10% condition be checked separately for the two samples when sampling without replacement?
  5. A student says, “The standard deviation is 0.06, so every sample difference will be within 0.06 of the mean.” Explain what the standard deviation actually describes.