Tutorials › AP Statistics › Pooled Proportion and Why It Is Used

Two-proportion hypothesis tests · Tutorial 523 of 1000

Pooled Proportion and Why It Is Used

See how the null hypothesis of equal proportions leads to one pooled estimate, and calculate it correctly from both groups’ success counts and sample sizes.

Intermediate 9 min read

What You'll Learn

  • Explain why equal population proportions under the null are represented by one common proportion.
  • Calculate the pooled sample proportion from the combined number of successes and observations.
  • Show why the pooled proportion is a weighted combination, not usually the simple average of sample proportions.
  • Distinguish the pooled estimate from the difference in sample proportions.
  • Recognize why pooling is used for a two-proportion hypothesis test but not a confidence interval.

Why a Two-Proportion Test Pools the Data

In Defining Both Parameters Clearly, you identified \(p_1\) and \(p_2\) as the true proportions for two groups and kept their order clear. In Hypotheses for Comparing Two Population Proportions, you saw that the null hypothesis for no difference can be written \(H_0:p_1=p_2\). That equality has an important consequence for the test calculation: if the null is true, both groups share one common population proportion.

The sample proportions, \(\hat{p}_1\) and \(\hat{p}_2\), will not usually be exactly equal. They come from different samples and can vary by chance. But when calculating a test under \(H_0\), we proceed as if both groups have the same true proportion. To estimate that shared proportion from the data, we combine the successes and observations from both samples. This estimate is called the pooled proportion.

Definition: The pooled proportion is the combined number of successes in both independent samples divided by the combined sample size. It estimates the common population proportion assumed by the null hypothesis \(H_0:p_1=p_2\).

Let \(x_1\) and \(x_2\) be the numbers of successes in the two samples, and let \(n_1\) and \(n_2\) be their sample sizes. Since \(\hat{p}_1=x_1/n_1\) and \(\hat{p}_2=x_2/n_2\), the pooled proportion is:

$$ \hat{p}_{\text{pool}}=\frac{x_1+x_2}{n_1+n_2} $$

The numerator counts every success across both groups. The denominator counts every observation across both groups. The pooled proportion is therefore a proportion for the combined data, used as an estimate of the common proportion under the null. It does not say the observed sample proportions are equal, and it does not erase the distinction between the two groups.

Why the Sample Sizes Matter

A pooled proportion is a weighted combination of the two sample proportions. A larger sample contributes more observations, so it has more influence on the combined estimate. If sample sizes are unequal, simply taking the average of \(\hat{p}_1\) and \(\hat{p}_2\) gives the two groups equal weight even though one group contributes more data. That is generally not the pooled proportion.

$$ \hat{p}_{\text{pool}} =\frac{n_1\hat{p}_1+n_2\hat{p}_2}{n_1+n_2} =\frac{x_1+x_2}{n_1+n_2} $$

This formula makes the weighting visible: multiply each sample proportion by its sample size to recover its success count, add those counts, and divide by the total sample size. If the two sample sizes happen to be equal, the pooled proportion is the ordinary average of the sample proportions. Otherwise, do not assume it is.

Worked Example: Estimating a Shared Transit-Use Proportion

Question: In a fictional survey, 38 of 80 sampled residents in one neighborhood used public transit at least once during the past week. In an independent sample from a second neighborhood, 27 of 60 residents did so. For the null hypothesis that the true neighborhood proportions are equal, calculate the pooled proportion.

Identify the counts: Group 1 has \(x_1=38\) successes out of \(n_1=80\), and Group 2 has \(x_2=27\) successes out of \(n_2=60\). The combined number of successes is \(38+27=65\), and the combined sample size is \(80+60=140\).

Calculate the pooled proportion:

$$ \hat{p}_{\text{pool}} =\frac{38+27}{80+60} =\frac{65}{140} \approx 0.4643 $$

Under the null hypothesis of equal population proportions, the pooled estimate of the common proportion who used public transit at least once that week is about \(0.4643\), or \(46.43\%\). As a check, the sample proportions are \(38/80=0.475\) and \(27/60=0.45\). Weighting them by their sample sizes gives \((80(0.475)+60(0.45))/140=(38+27)/140\approx0.4643\), the same result.

The observed proportions, \(0.475\) and \(0.45\), are not identical. Pooling does not claim they are. It estimates the single common proportion that the null hypothesis says generated outcomes in both groups.

Pooling Belongs to the Null Model

The pooled proportion is used in a two-proportion hypothesis test because the test calculations are made under the assumption that \(H_0\) is true. If \(p_1=p_2\), both groups have the same underlying success probability, so their observed successes can be combined to estimate it. This shared estimate is then used in the test’s null model.

The pooled proportion is not the estimate of the difference \(p_1-p_2\). For the observed samples, the point estimate of that difference remains \(\hat{p}_1-\hat{p}_2\), in the group order defined for the problem. The pooled proportion answers a different question: what common success proportion is estimated when we impose the equality stated by \(H_0\)?

Pooling is specific to the null hypothesis of equal proportions. A two-proportion confidence interval does not assume \(p_1=p_2\); it estimates the difference that may exist. As explained in Common Errors With Two-Proportion Intervals, the interval uses separate sample proportions in its standard error rather than pooling. Do not carry the pooled calculation over to an interval just because both methods compare two proportions.

Key distinction: In a two-proportion hypothesis test with \(H_0:p_1=p_2\), use the pooled proportion to estimate the common success proportion under the null. The sample difference \(\hat{p}_1-\hat{p}_2\) still describes the observed group difference. A two-proportion confidence interval uses a different approach and does not pool.

Worked Example: Combining Results From Unequal Samples

Question: A fictional sports program asks whether two training groups have the same proportion of participants who complete a set of drills. In Group 1, 54 of 90 participants complete the drills; in Group 2, 39 of 60 do. Calculate the pooled proportion for the null hypothesis of equal population proportions.

Combine successes and sample sizes: There are \(54+39=93\) successes among \(90+60=150\) participants. Thus,

$$ \hat{p}_{\text{pool}} =\frac{54+39}{90+60} =\frac{93}{150} =0.62 $$

The pooled estimate of the common completion proportion under the null is \(0.62\), or \(62\%\). To check the weighting, the sample proportions are \(54/90=0.60\) and \(39/60=0.65\). Their weighted average is \((90(0.60)+60(0.65))/150=(54+39)/150=0.62\).

The simple average of the two sample proportions would be \((0.60+0.65)/2=0.625\), which is not the pooled proportion. Group 1 has 90 observations and Group 2 has 60, so the combined estimate must reflect those different sample sizes. The pooled value is closer to \(0.60\), the proportion from the larger sample.

Worked Example: Avoiding an Unweighted Average

Question: In a fictional consumer survey, 18 of 30 customers at one type of store scan a product’s information code, compared with 42 of 120 customers at another type of store. Calculate the pooled estimate for a test of equal population proportions.

Count the combined data: The two samples contain \(18+42=60\) successes and \(30+120=150\) customers. Therefore,

$$ \hat{p}_{\text{pool}} =\frac{18+42}{30+120} =\frac{60}{150} =0.40 $$

The pooled proportion is \(0.40\), or \(40\%\). The sample proportions are \(18/30=0.60\) and \(42/120=0.35\). The unweighted average is \((0.60+0.35)/2=0.475\), which is quite different from \(0.40\). It gives the 30-customer sample the same influence as the 120-customer sample, despite the latter having four times as many observations.

The weighted calculation confirms the pooled result: \((30(0.60)+120(0.35))/150=(18+42)/150=0.40\). The observed difference is \(0.60-0.35=0.25\), but the pooled estimate is not that difference. The value \(0.40\) estimates the shared success proportion assumed by the null hypothesis.

What Pooling Does—and Does Not—Mean

Combining counts is a calculation for a particular model, not a claim that the two groups are interchangeable in every respect. The groups remain separately identified, their sample sizes remain \(n_1\) and \(n_2\), and the difference in their sample proportions remains available to compare with the null model. Pooling uses the equality in \(H_0\) to estimate one common success probability.

A pooled proportion can also be used to understand the Large Counts condition for a two-proportion \(z\)-test. Under the equal-proportions null model, the expected numbers of successes in the two groups are \(n_1\hat{p}_{\text{pool}}\) and \(n_2\hat{p}_{\text{pool}}\); the expected numbers of failures are \(n_1(1-\hat{p}_{\text{pool}})\) and \(n_2(1-\hat{p}_{\text{pool}})\). The test’s Large Counts check considers whether these expected counts are at least 10 in each group. This is distinct from checking observed successes and failures for a two-proportion confidence interval.

For instance, in the transit example, the pooled estimate is about \(0.4643\). Under the null model, the expected success and failure counts are about \(80(0.4643)=37.14\) and \(80(1-0.4643)=42.86\) for Group 1, and \(60(0.4643)=27.86\) and \(60(1-0.4643)=32.14\) for Group 2. These are expected counts under the null, not the actual observed counts. Each exceeds 10. The randomization or random-sampling and independence conditions still need to be considered separately, as covered in earlier tutorials.

Common Mistakes and AP Exam Tips

  • Averaging the two proportions without weighting: Unless the sample sizes are equal, \((\hat{p}_1+\hat{p}_2)/2\) generally is not pooled. Add success counts and divide by the combined sample size.
  • Adding sample sizes but not successes: The numerator must include successes from both groups; the denominator must include all observations from both groups.
  • Using the pooled value as the observed difference: The pooled proportion estimates a common success probability under \(H_0\). The observed difference is \(\hat{p}_1-\hat{p}_2\).
  • Assuming the observed proportions must match: \(H_0:p_1=p_2\) concerns the population proportions. Sample proportions can differ because of sampling variability.
  • Pooling for a confidence interval: The interval does not impose the equality assumption in the null hypothesis. Use the separate sample proportions for a two-proportion interval.
  • Rounding too early: Keep the fraction or several decimal places during calculations, then round the reported pooled estimate. This helps avoid small discrepancies later.
AP Exam Tip: Write the pooled formula with the counts visible: combined successes divided by combined sample size. Then state what the result estimates— the common population proportion assumed by \(H_0\). Do not describe it as proof that the population proportions are equal.

Key Takeaway

The equality in \(H_0:p_1=p_2\) gives a two-proportion test one common success proportion to estimate. Pool the two groups’ success counts and divide by their combined sample size; this automatically weights each group according to how many observations it contributes.

Key takeaway: For a two-proportion test of equal population proportions, calculate \(\hat{p}_{\text{pool}}=(x_1+x_2)/(n_1+n_2)\). It estimates the shared proportion under the null, not the observed difference and not a confidence-interval estimate.

Check Your Understanding

For each question, focus on what the pooled proportion estimates and how to calculate it.

  1. In one sample, 24 of 40 people report using a refillable bottle; in a second sample, 35 of 70 do. Calculate the pooled proportion for a test of equal population proportions.
  2. Why does \(H_0:p_1=p_2\) motivate combining successes from both samples?
  3. Two samples have proportions \(0.30\) and \(0.50\), with sample sizes 20 and 100. Explain why their simple average is not the pooled proportion.
  4. In a two-proportion test, what does \(\hat{p}_{\text{pool}}\) estimate, and what does \(\hat{p}_1-\hat{p}_2\) describe?
  5. Should a pooled proportion be used in a two-proportion confidence interval? Explain briefly.