Why Group Order Matters
In Intervals Entirely Above or Below Zero, you learned to use the sign of a confidence interval for \(p_1-p_2\) to identify which group has the higher population proportion. The sign only makes sense when you know which group is first. This tutorial focuses on choosing that order and keeping it consistent from the parameter definition through the final interpretation.
There is no statistical rule that makes one group automatically Group 1. The question may specify an order, or you may choose one that makes the comparison easy to describe. Either way, state the order explicitly. If the question asks for Group A minus Group B, define Group A as Group 1. If it asks for Group B minus Group A, define Group B as Group 1—even if a table happens to list Group A first.
Changing the order changes the signs of the difference and both interval endpoints. It does not change the underlying comparison. For example, if an interval for \(p_1-p_2\) is entirely positive, the equivalent interval for \(p_2-p_1\) is entirely negative. Both support the same conclusion: Group 1 has the higher population proportion.
Choose the Order, Then Keep It
Before calculating, identify the characteristic that counts as a success and name the two groups in the order required by the question. Then define \(p_1\) and \(p_2\) using that order. In calculations, pair each group’s success count with its own sample size: \(\hat{p}_1=x_1/n_1\) and \(\hat{p}_2=x_2/n_2\). Do not switch labels partway through.
If the question does not specify an order, choose one that makes the comparison natural—for example, a new option minus a standard option, or a named location minus another named location. Either choice is valid as long as the order is stated and used consistently. Reversing the groups is also valid, but the reported parameter and interpretation must reverse with them.
The endpoint order matters. If the original interval is \((0.03,0.25)\), the reverse interval is \((-0.25,-0.03)\), not \((-0.03,-0.25)\). The point estimate changes from \(\hat{p}_1-\hat{p}_2\) to its negative, and the standard error and margin of error stay the same. Thus the interval shifts to the opposite side of zero while preserving its width.
This is useful when a question gives a confidence interval in one order but asks for the comparison in the other order. You do not need to construct a new interval from scratch. Transform the endpoints, then interpret the new parameter in context. The direction of the conclusion should agree in both versions.
Worked Examples
Worked Example: Choose the Order Requested
Setting: Imagine independent random samples of customers from two meal-delivery services. A success is a customer who says the service’s delivery time is dependable. Service North has 84 successes among 120 sampled customers; Service South has 66 among 120. Assume each service has at least 1,200 customers in the population of interest. Construct and interpret a 95% confidence interval for the proportion for North minus the proportion for South.
State: Let \(p_N\) and \(p_S\) be the true proportions of customers using Services North and South, respectively, who say delivery time is dependable. Since the requested order is North minus South, define Group 1 as North and Group 2 as South. The parameter is \(p_N-p_S=p_1-p_2\).
Plan: The data come from independent random samples, supporting the Random condition and independence between groups. Each sample of 120 is at most 10% of its population because each service has at least 1,200 customers. For the Large Counts condition, North has 84 successes and \(120-84=36\) failures; South has 66 successes and \(120-66=54\) failures. All four counts are at least 10, so the condition is met.
Do: The sample proportions and their difference are:
Use the unpooled standard error for a two-proportion confidence interval:
For 95% confidence, \(z^*\approx1.959964\). The margin of error is \(1.959964(0.061745)\approx0.12102\). Therefore:
Conclude: We are 95% confident that the true proportion of Service North customers who say delivery time is dependable minus the true proportion of Service South customers is between about 0.029 and 0.271. The interval is entirely positive, supporting the conclusion that the population proportion is higher for Service North. The plausible difference is about 2.9 to 27.1 percentage points.
Worked Example: Reverse the Group Order
Setting: Imagine independent random samples of 100 households from each of two recycling districts. A success is a household that separates food scraps for composting. District Cedar has 52 successes and District Maple has 68. Assume each district has at least 1,000 households. Find a 90% confidence interval for Cedar minus Maple, then express the same result as Maple minus Cedar.
State: Let \(p_C\) and \(p_M\) be the true proportions of households in Cedar and Maple, respectively, that separate food scraps for composting. First estimate \(p_C-p_M\), so Cedar is Group 1 and Maple is Group 2.
Plan: The independent random samples support the Random condition and independence between groups. Each sample of 100 is at most 10% of its population because each district has at least 1,000 households. Cedar has 52 successes and \(100-52=48\) failures; Maple has 68 successes and \(100-68=32\) failures. All four counts are at least 10, meeting the Large Counts condition.
Do: The point estimate is:
The standard error and interval are:
For 90% confidence, \(z^*\approx1.644854\). The margin of error is \(1.644854(0.068352)\approx0.11243\), so the interval is:
To reverse the order, negate the endpoints and switch their positions:
Conclude: We are 90% confident that the true composting proportion in Cedar minus the true proportion in Maple is between about \(-0.272\) and \(-0.048\). Since this interval is entirely negative, it supports the conclusion that Maple has the higher population proportion. In the reverse order, the interval is entirely positive and gives the same conclusion: Maple’s proportion is higher, by about 4.8 to 27.2 percentage points.
Worked Example: Let the Question Set Group 1
Setting: Imagine independent random samples of households from two municipal service zones. A success is a household that reports using a neighborhood compost drop-off site at least once a month. A sample from Zone East has 108 successes among 200 households, and a sample from Zone West has 132 among 200. Assume each zone has at least 2,000 households. The question asks for West minus East.
State: Let \(p_W\) and \(p_E\) be the true proportions of households in West and East, respectively, that use a compost drop-off site monthly. Although East was listed first in the setting, the requested parameter is \(p_W-p_E\). Therefore, West is Group 1 and East is Group 2.
Plan: These are independent random samples, supporting the Random condition and independence between groups. Each sample of 200 is at most 10% of its population because each zone has at least 2,000 households. West has 132 successes and \(200-132=68\) failures; East has 108 successes and \(200-108=92\) failures. Each count is at least 10, so the Large Counts condition is satisfied.
Do: Calculate the requested estimate in the requested order:
The standard error is:
For 95% confidence, \(z^*\approx1.959964\), giving a margin of error of \(1.959964(0.048621)\approx0.09530\). Thus:
Conclude: We are 95% confident that the true proportion of households in West that use a compost drop-off site monthly minus the true proportion in East is between about 0.025 and 0.215. The entire interval is positive, supporting the conclusion that West has the higher population proportion. If the question instead asked for East minus West, the interval would be \((-0.21530,-0.02470)\); that negative interval would still indicate that West is higher.
Common Mistakes and AP Exam Tips
- Assuming the first group listed must be Group 1: The parameter order comes from the question or your clearly stated choice, not from the order in a table. If the requested quantity is West minus East, calculate West first.
- Changing order in the calculation but not in the interpretation: A positive interval for \(p_1-p_2\) favors Group 1. A positive interval for \(p_2-p_1\) favors Group 2 relative to the original labels. Define the groups before interpreting the sign.
- Negating endpoints without rearranging them: Reversing \((L,U)\) gives \((-U,-L)\). Keeping the endpoints in the wrong order would no longer be a valid lower-to-upper interval.
- Thinking the reversed interval changes the evidence: The numerical signs change, but the interval width and the conclusion about which group has the higher proportion do not. The change is a matter of expressing the same comparison in the opposite order.
- Using vague wording: “The difference is positive” does not identify the groups or characteristic. A full-credit interpretation names the population proportions, says which is subtracted from which, and translates the sign into which group has the higher proportion.
- Confusing percentage points with percent change: A difference of 0.12 is 12 percentage points. It is not automatically a 12% increase. Use the units that describe the difference between proportions.
Key Takeaway
You may choose Group 1 to match the question or to make the comparison easier to communicate, but you must declare the order and preserve it throughout. Reversing the order negates the difference and reverses the endpoints, yet it does not change which group has the higher population proportion.
Check Your Understanding
Assume all intervals are valid confidence intervals. Show how the group order affects the signs and write each conclusion in context when possible.
- A question asks for Group B minus Group A. Which group should be defined as Group 1?
- An interval for \(p_1-p_2\) is \((0.04,0.19)\). Give the equivalent interval for \(p_2-p_1\).
- An interval for North minus South is \((-0.16,-0.03)\). Which group has the higher population proportion according to the interval?
- Why is \((-0.03,-0.16)\) not a correctly ordered interval when reversing \((0.03,0.16)\)?
- If two equivalent intervals have opposite signs, why can they still support the same conclusion about which group has the higher proportion?