From Theoretical Spread to an Estimated Standard Error
In Sampling Distribution of the Difference in Sample Proportions, you saw that the standard deviation of \(\hat{p}_1-\hat{p}_2\) depends on the true population proportions \(p_1\) and \(p_2\). In practice, those population proportions are unknown. For a confidence interval, we estimate the spread using the sample proportions instead.
The result is the standard error of \(\hat{p}_1-\hat{p}_2\). It estimates how much the difference in sample proportions would typically vary from sample to sample, for samples of these sizes and taken in the same way. Keep the group order fixed: the statistic is Group 1 minus Group 2.
The Formula and Its Parts
For two independent samples, find each sample proportion by dividing that group’s number of successes by its sample size. Then use the two sample proportions in the formula below. The hats are important: this is an estimated standard deviation based on sample data, not the theoretical standard deviation that uses \(p_1\) and \(p_2\).
Each fraction estimates the variance contribution from one group’s sample proportion. Add those two contributions, then take the square root. In particular, do not subtract one group’s contribution from the other: the two sources of sampling variation add.
The standard error is nonnegative, even when \(\hat{p}_1-\hat{p}_2\) is negative. It describes estimated spread, not the direction of the observed difference. The point estimate \(\hat{p}_1-\hat{p}_2\) carries the direction; the standard error measures the typical sampling variation around that estimate.
The formula keeps the groups separate. Do not combine the two success counts and sample sizes to make one pooled proportion for this calculation. Each group has its own estimated proportion and its own variance contribution.
Standard error is expressed in proportion units. For example, an SE of \(0.05\) is equivalent to 5 percentage points. It is not a guaranteed maximum error, and it is not the standard deviation of the individual yes-or-no outcomes. It estimates the variability of the statistic \(\hat{p}_1-\hat{p}_2\) across repeated pairs of samples.
A Reliable Calculation Routine
A careful calculation makes it easier to spot errors and explain what the result means. Work through the groups in order, keep enough digits during the arithmetic, and round the final standard error only after taking the square root.
- Record \(x_1,n_1,x_2,n_2\), and confirm which group is Group 1.
- Calculate \(\hat{p}_1=x_1/n_1\) and \(\hat{p}_2=x_2/n_2\).
- Calculate each contribution, \(\hat{p}_1(1-\hat{p}_1)/n_1\) and \(\hat{p}_2(1-\hat{p}_2)/n_2\).
- Add the contributions and take the square root.
- State the result in proportion units and, if helpful, convert it to percentage points.
This routine is about calculating the standard error; it does not by itself establish that a Normal-based confidence interval is appropriate. As in the earlier tutorial on the sampling distribution, the study design and relevant conditions matter for inference. The calculation examples below focus on the standard error and state the sampling setup so that the groups can be treated as independent.
Worked Examples
Worked Example: Comparing Two School Survey Proportions
Setting: Imagine independent random samples of students from two large school districts. In District 1, 64 of 100 sampled students say they usually bring a reusable water bottle. In District 2, 52 of 100 say they do. Suppose each district has at least 10 times its sample size in students. Find the standard error for the difference, using District 1 minus District 2.
Find the sample proportions:
Substitute into the formula: The first group contributes \(0.64(1-0.64)/100\), and the second contributes \(0.52(1-0.52)/100\).
Check the arithmetic: The two products are \(0.64(0.36)=0.2304\) and \(0.52(0.48)=0.2496\). Dividing each by 100 gives \(0.002304\) and \(0.002496\); their sum is \(0.004800\). Squaring \(0.0693\) gives approximately \(0.004802\), consistent with the unrounded square root of \(0.004800\).
The estimated standard error is about \(0.0693\), or 6.93 percentage points. It estimates the typical sample-to-sample variation in the difference in reusable-bottle proportions for independent samples of these sizes. It does not change if the group subtraction order is reversed; only the point estimate’s sign would change.
Worked Example: Unequal Sample Sizes in an Environmental Survey
Setting: Imagine independent random samples from two large communities. In Community 1, 96 of 160 sampled households report composting food scraps. In Community 2, 54 of 120 sampled households report doing so. Suppose both population sizes are at least 10 times their respective sample sizes. Calculate the standard error for Community 1 minus Community 2.
Calculate the sample proportions:
Calculate the two contributions: The sample sizes differ, so divide each group’s variance contribution by its own sample size.
Check the arithmetic: The first contribution is \(0.24/160=0.001500\). The second is \(0.2475/120=0.0020625\). They add to \(0.0035625\); squaring \(0.0597\) gives approximately \(0.003564\), which agrees with the square root rounded to four decimal places.
The estimated standard error is about \(0.0597\), or 5.97 percentage points. Although Community 1 has a larger sample, both terms matter. The second contribution is larger here, so it contributes more to the estimated spread. Always use \(n_1\) with Group 1’s proportion and \(n_2\) with Group 2’s proportion.
Worked Example: A Larger Standard Error with Smaller Samples
Setting: Imagine independent random samples from two large groups of online shoppers. In Group 1, 27 of 45 sampled shoppers say they would choose a same-day delivery option. In Group 2, 21 of 60 say they would. Suppose the groups were sampled independently and each population is at least 10 times its sample size. Find the standard error for Group 1 minus Group 2.
Find the sample proportions:
Substitute and calculate:
Check the arithmetic: The first contribution is \(0.24/45\approx0.0053333\). The second is \(0.2275/60\approx0.0037917\). Adding the unrounded fractions gives \(0.009125\), whose square root is approximately \(0.095525\). Rounded to four decimal places, the standard error is \(0.0955\).
The standard error is about \(0.0955\), or 9.55 percentage points. That is larger than the standard errors in the earlier examples, in part because these samples are smaller. A larger SE indicates greater estimated sample-to-sample variability; it does not say that the observed sample difference is wrong.
Common Mistakes and AP Exam Tip
- Using the wrong proportions: Use \(\hat{p}_1\) and \(\hat{p}_2\) for this standard error, not unknown \(p_1\) and \(p_2\), and not a null value from a hypothesis test.
- Pooling the samples: Do not combine the successes and sample sizes into one proportion. The formula estimates a separate variance contribution for each group.
- Subtracting the variance contributions: Add the two terms under the square root. The samples contribute separate sources of variation to the difference.
- Pairing the wrong sample size with a group: Keep each group’s proportion and sample size together. Reversing or mixing subscripts can change the result.
- Taking the square root too early: First calculate and add the two contributions, then take one square root. Keep extra digits until the final rounding.
- Calling the SE the observed difference: The observed difference is \(\hat{p}_1-\hat{p}_2\). The SE is a nonnegative estimate of that statistic’s sampling variability.
- Overstating what the number guarantees: An SE is not a maximum possible error and does not guarantee how close one sample result is to the population difference.
Key Takeaway
The standard error for a difference in sample proportions estimates the sampling distribution’s spread using the observed proportions. Calculate each group’s proportion from its own success count and sample size, add the two separate variance contributions, and take the square root. Interpret the result as estimated sample-to-sample variability, not as the observed difference or a guaranteed error.
Check Your Understanding
For each question, show the sample proportions and the separate variance contributions before giving the standard error.
- In two independent samples, 72 of 120 Group 1 participants and 45 of 90 Group 2 participants meet a stated criterion. Calculate the standard error for \(\hat{p}_1-\hat{p}_2\).
- Why does the standard-error formula use the sample proportions rather than the unknown population proportions?
- A student adds the two variance contributions, then reports that sum as the standard error. What calculation is missing?
- If the observed difference \(\hat{p}_1-\hat{p}_2\) is negative, must its standard error also be negative? Explain.
- Explain why a standard error of \(0.04\) does not guarantee that every sample difference is within \(0.04\) of the population difference.