From Two Sample Proportions to an Interval
In Large Counts for Each Group in Two-Proportion Intervals, you learned to check whether each group has enough successes and failures to support a Normal-based interval. Now we will use the sample proportions and those conditions to construct the full interval by hand.
The focus example compares 48 successes in a sample of 100 with 35 successes in another sample of 100. The difference between the sample proportions is the point estimate for the difference between the two population proportions. But a point estimate alone does not show how much the estimate might vary from sample to sample. A confidence interval adds a margin of error to describe that uncertainty.
The Hand Calculation
As established in Comparing Two Proportions: Setting and Notation, keep the group order consistent. Here, Group 1 is the first sample and Group 2 is the second, so the interval estimates \(p_1-p_2\). The sample proportions are calculated separately, using each group’s successes divided by its own sample size.
The standard error for the interval estimates the variability of the difference in sample proportions. As explained in Standard Error for a Difference in Proportions, it uses each sample’s own proportion. Do not combine the samples into a pooled proportion for this interval.
Choose a critical value \(z^*\) that matches the confidence level. For a 95% confidence interval, \(z^*\) is approximately 1.96. The margin of error is \(z^*\) times the standard error, and the interval is the point estimate plus or minus that margin of error.
Worked Examples
Worked Example: The 48/100 Versus 35/100 Example
Setting: Imagine independent random samples of residents from two towns. A survey asks whether each resident uses public transit to travel to work. In Town 1, 48 of 100 sampled residents say yes; in Town 2, 35 of 100 say yes. Assume each town has at least 1,000 residents. Construct and interpret a 95% confidence interval for the difference in the true proportions, in Town 1 minus Town 2.
State: Let \(p_1\) be the true proportion of Town 1 residents who use public transit to travel to work, and let \(p_2\) be the corresponding proportion for Town 2. We want an interval estimating \(p_1-p_2\).
Plan: The two samples are stated to be independent random samples, supporting the Random condition and independence between groups. Within each sample, the 10% condition is met because 100 is no more than 10% of a town population of at least 1,000. The Large Counts condition is also met: Town 1 has 48 successes and \(100-48=52\) failures; Town 2 has 35 successes and \(100-35=65\) failures. All four counts are at least 10, so we can use a two-proportion z-interval.
Do: First calculate each sample proportion and their difference:
Next calculate the interval’s standard error using the separate sample proportions:
For 95% confidence, use \(z^*=1.96\). The margin of error is:
Subtract and add the margin of error to the point estimate:
Conclude: We are 95% confident that the proportion of Town 1 residents who use public transit to travel to work is between about 0.005 lower and 0.265 higher than the corresponding proportion for Town 2. In percentage-point terms, the plausible difference ranges from about 0.54 percentage points lower to 26.54 percentage points higher. Since zero is in the interval, the data are compatible with no difference between the population proportions as well as with differences in either direction.
Worked Example: A 90% Interval for Two Service Groups
Setting: Imagine independent random samples of customers from two online services. Customers are asked whether they have enabled a particular account notification. In Service 1, 62 of 100 sampled customers have enabled it; in Service 2, 48 of 100 have enabled it. Assume both customer populations have at least 1,000 people. Construct a 90% confidence interval for \(p_1-p_2\), where the proportions refer to customers who enabled the notification in Services 1 and 2, respectively.
Plan and conditions: The samples are described as independent random samples, and each sample is no more than 10% of its population. Service 1 has 62 successes and 38 failures; Service 2 has 48 successes and 52 failures. Each count is at least 10, so the Large Counts condition is met.
Do: The point estimate is \(62/100-48/100=0.62-0.48=0.14\). Calculate the standard error:
For a 90% confidence interval, use \(z^*\approx1.645\). The margin of error is \(1.645(0.06966)\approx0.11458\). Therefore:
Conclude: We are 90% confident that the true proportion of Service 1 customers who enabled the notification is about 0.025 to 0.255 higher than the true proportion for Service 2. This is an interval for a difference in population proportions, not two separate intervals for the individual proportions.
Worked Example: An Interval That Includes Zero
Setting: Imagine independent random samples of 80 and 75 gardeners from two regions. In Region 1, 40 sampled gardeners use a particular watering method; in Region 2, 33 do. Assume each region has at least 800 gardeners. Construct a 95% confidence interval for the difference \(p_1-p_2\), with Region 1 first.
Plan and conditions: The samples are independent random samples. Each sample is no more than 10% of its population. The success and failure counts are 40 and 40 in Region 1, and 33 and 42 in Region 2; all are at least 10. Thus the standard two-proportion interval is appropriate.
Do: The sample proportions are \(40/80=0.50\) and \(33/75=0.44\), so the point estimate is \(0.50-0.44=0.06\). The standard error is:
Using \(z^*=1.96\), the margin of error is \(1.96(0.08006)\approx0.15693\). The interval is:
Conclude: We are 95% confident that the proportion using the watering method in Region 1 is between about 0.097 lower and 0.217 higher than the proportion in Region 2. Because zero lies within the interval, these data do not pin down the direction of the population difference at this confidence level.
How to Interpret the Result
An interval’s endpoints describe plausible values for the difference \(p_1-p_2\), using the group order stated in the question. A positive value means the proportion is higher in Group 1; a negative value means it is lower in Group 1. An interval containing zero includes “no difference” among the plausible values. It does not prove that the population proportions are equal.
The confidence level describes the long-run performance of the method. If we repeatedly took samples in the same way and constructed an interval each time, about 95% of the intervals produced by a 95% method would capture the true difference \(p_1-p_2\). It is not correct to say that there is a 95% probability that this particular fixed parameter is inside the interval.
The order of subtraction matters. Reversing the groups changes \(p_1-p_2\) to \(p_2-p_1\), so the interval’s endpoints change sign and switch order. Clearly name the groups and preserve that order from the parameter definition through the final interpretation.
Common Mistakes and AP Exam Tip
- Using a pooled proportion: This interval estimates a difference using the two observed sample proportions separately. A pooled proportion does not belong in its standard error.
- Using the wrong sample size: In each standard-error term, pair a group’s own sample proportion with that group’s own sample size.
- Reversing the subtraction: If the parameter is \(p_1-p_2\), the point estimate must be \(\hat{p}_1-\hat{p}_2\), and the conclusion must keep Group 1 first.
- Giving only the point estimate: The sample difference \(0.13\) is not the confidence interval. Show the standard error, critical value, margin of error, and both endpoints.
- Misreading zero in the interval: An interval that includes zero does not show that the proportions are equal. It means zero remains one plausible value for their difference.
- Claiming a probability about a fixed parameter: A full-credit interpretation describes the long-run capture rate of the method, then states the interval’s plausible values in context.
Key Takeaway
To construct a two-proportion z-interval by hand, check the design and Large Counts conditions, calculate \(\hat{p}_1-\hat{p}_2\), find the standard error using each group’s own sample proportion, and add and subtract \(z^*SE\). Interpret the endpoints as plausible values for the difference between the two population proportions in context.
Check Your Understanding
Use the two-proportion interval method to answer each question. Show your calculations where appropriate.
- In the main example, what is the point estimate for \(p_1-p_2\), and what does its sign say about the sample proportions?
- Why does the interval standard error use \(\hat{p}_1\) and \(\hat{p}_2\) separately rather than a pooled proportion?
- If a 95% interval for \(p_1-p_2\) is \((-0.0054,\ 0.2654)\), what does zero’s presence in the interval tell you?
- In the service example, identify the critical value and margin of error used to form the 90% interval.
- How would the endpoints and interpretation change if the group order were reversed?