From a Pooled Proportion to a Pooled Standard Error
In Pooled Proportion and Why It Is Used, you learned to combine successes and sample sizes to estimate the common success proportion assumed by \(H_0:p_1=p_2\). The next calculation uses that estimate to measure the variability expected in the difference between two sample proportions when the null hypothesis is true.
The resulting quantity is called the pooled standard error. It is the estimated standard deviation of \(\hat{p}_1-\hat{p}_2\) under the null model of equal population proportions. It uses the same pooled estimate for both groups, but it still accounts for each group’s own sample size. This tutorial focuses on calculating that standard error; it is one part of a two-proportion hypothesis test, not the whole test.
Let \(x_1\) and \(x_2\) be the success counts in the two samples, with sample sizes \(n_1\) and \(n_2\). The pooled proportion is \(\hat{p}_c=(x_1+x_2)/(n_1+n_2)\). This is the same quantity called \(\hat{p}_{\text{pool}}\) in the previous tutorial; the subscript \(c\) emphasizes that it estimates the common proportion assumed under the null.
The expression has two parts. The quantity \(\hat{p}_c(1-\hat{p}_c)\) describes the estimated variation in a single success-or-failure outcome under the common-proportion model. The sum \(1/n_1+1/n_2\) accounts for the contribution of both independent samples. Take the square root at the end: the result is in proportion units, just like \(\hat{p}_1-\hat{p}_2\).
Worked Calculation
Worked Example: Comparing Two Transit Surveys
Question: In a fictional study, independent random samples are taken from two large neighborhoods. In Neighborhood 1, 42 of 70 sampled residents used public transit at least once last week. In Neighborhood 2, 30 of 60 did. Calculate the pooled standard error for testing whether the population proportions are equal.
Find the pooled proportion: There are \(42+30=72\) successes among \(70+60=130\) residents. Therefore,
Substitute into the standard-error formula: Use \(n_1=70\), \(n_2=60\), and retain the unrounded fraction for the pooled proportion during the calculation.
The pooled standard error is about \(0.0875\), or 8.75 percentage points. Under the equal-proportions null model, differences between the sample proportions would typically vary by about 0.0875 from sample to sample. This is an estimate of variability, not the observed difference between these neighborhoods.
Check the calculation another way: The product \(\hat{p}_c(1-\hat{p}_c)\) is about \(0.2471\), and \(1/70+1/60=13/420\approx0.0310\). Their product is about \(0.007648\); its square root is again about \(0.0875\).
For a test, the Large Counts condition is checked using expected counts under the null model. Here, the expected successes are \(70(72/130)\approx38.77\) and \(60(72/130)\approx33.23\); the expected failures are \(70(58/130)\approx31.23\) and \(60(58/130)\approx26.77\). All four expected counts are at least 10. The scenario specifies independent random samples, and the neighborhoods are large enough for each sample to be less than 10% of its population, so the random, independence, and 10% conditions are supported as well.
Why the Formula Uses the Pooled Value
The pooled standard error belongs to the test’s null model. Under \(H_0:p_1=p_2\), both groups share one population success proportion. The pooled proportion estimates that common value, and the formula uses it in both groups’ contributions to the variability.
The sample sizes do not disappear when the outcomes are pooled. A group with a smaller sample size contributes more variability because its reciprocal sample size, \(1/n_i\), is larger. Increasing either sample size reduces its contribution to the sum in the formula, all else held constant.
This is different from the standard error for a two-proportion confidence interval, which was covered in Standard Error for a Difference in Proportions. A confidence interval does not impose \(p_1=p_2\), so it uses \(\hat{p}_1\) and \(\hat{p}_2\) separately. The pooled version is for a hypothesis test whose null hypothesis asserts equal population proportions. Match the standard-error formula to the inference question.
Worked Example: Unequal Sample Sizes
Question: In a fictional community survey, 25 of 50 households in one area report composting food scraps, while 45 of 100 households in a second area do. Find the pooled standard error for a test of equal population proportions.
Calculate the pooled proportion: The combined success count is \(25+45=70\), and the combined sample size is \(50+100=150\).
Calculate the pooled standard error:
The pooled standard error is about \(0.0864\), or 8.64 percentage points. As a check, \((7/15)(8/15)=56/225\approx0.2489\), and \(1/50+1/100=0.03\). Multiplying gives about \(0.007467\), whose square root is \(0.0864\).
Notice that the two sample sizes are not equal. The formula uses \(1/50+1/100\), not twice either reciprocal and not an average sample size. Each group contributes according to its own sample size. The pooled proportion is also not the simple average of \(25/50=0.50\) and \(45/100=0.45\); it is the combined success count divided by the combined sample size.
More Practice With the Calculation
Worked Example: A School Technology Survey
Question: Two independent random samples of students at large schools are asked whether they use a study-planning app each week. At School A, 16 of 40 students say yes. At School B, 54 of 120 say yes. Calculate the pooled standard error for testing whether the population proportions are equal.
Pool the successes: There are \(16+54=70\) students who say yes in a total of \(40+120=160\) sampled students.
Substitute the values:
The pooled standard error is about \(0.0906\), or 9.06 percentage points. To verify, the pooled success-failure product is \(0.4375(0.5625)=0.24609375\), and the reciprocal sample sizes sum to \(1/40+1/120=1/30\). Their product is \(0.008203125\), whose square root rounds to \(0.0906\).
The pooled estimate does not make the observed sample proportions equal: they are \(16/40=0.40\) and \(54/120=0.45\). It estimates the common population proportion assumed by the null hypothesis. The standard error describes how much the difference between sample proportions would vary under that null model.
Common Mistakes and AP Exam Tips
- Using separate sample proportions in the test formula: For the pooled standard error under \(H_0:p_1=p_2\), use \(\hat{p}_c\) in both places. Separate sample proportions belong in the unpooled standard error for a confidence interval.
- Pooling the sample sizes but not the successes: Calculate \(\hat{p}_c\) as combined successes divided by combined sample size. Do not average the two sample proportions unless the sample sizes are equal.
- Forgetting one group’s contribution: Include both \(1/n_1\) and \(1/n_2\). Omitting either makes the calculated standard error too small.
- Taking the square root too early—or not at all: Multiply the factors inside the radical first, then take the square root. The value inside the radical is a variance-like quantity, not the standard error.
- Confusing standard error with observed difference: The standard error measures estimated sampling variability. The observed difference is \(\hat{p}_1-\hat{p}_2\), with the group order stated in the problem.
- Rounding intermediate values too much: Keep the pooled fraction or several decimal places through the calculation. Round the final standard error to a sensible number of decimal places.
Key Takeaway
The pooled standard error quantifies the expected sampling variability of the difference in sample proportions under the equal-proportions null hypothesis. Calculate the common pooled estimate first, then let each sample size contribute through its reciprocal.
Check Your Understanding
Use the pooled standard-error formula and keep the sample sizes paired with their own groups.
- In two independent samples, 20 of 40 people in Group 1 and 36 of 80 people in Group 2 report carrying a reusable cup. Calculate \(\hat{p}_c\) and the pooled standard error.
- Explain why a two-proportion test with \(H_0:p_1=p_2\) uses the same pooled estimate for both groups.
- If \(n_1\) increases while the pooled proportion and \(n_2\) stay fixed, what happens to the contribution \(1/n_1\), and how does that affect the pooled standard error?
- For a test, the pooled proportion is \(0.40\), with sample sizes \(n_1=50\) and \(n_2=100\). Calculate the pooled standard error.
- Why is the pooled standard error not the one used for a two-proportion confidence interval?