When Every Plausible Difference Has the Same Sign
In What It Means When the Interval Contains Zero, you learned that zero in an interval for \(p_1-p_2\) means no difference is one plausible value, although nonzero differences may also be plausible. This tutorial considers the contrasting situation: zero is not in the interval, and both endpoints are on the same side of zero.
The sign of an interval’s values identifies the direction of the difference, provided you keep the group order straight. A positive value of \(p_1-p_2\) means \(p_1\) is greater than \(p_2\). A negative value means \(p_1\) is less than \(p_2\). If an entire confidence interval is positive, all the differences it identifies as plausible favor Group 1; if the entire interval is negative, they favor Group 2.
This is a conclusion about the direction of the difference in the population proportions, supported by the interval—not just a description of which sample proportion was larger. As in Interpreting a Confidence Interval for \(p_1-p_2\), name the characteristic and both groups, and state the subtraction order. Do not say the interval proves the proportions differ or gives the probability that one group is higher.
Read the Sign Before Naming the Higher Group
First check the parameter: the interval estimates \(p_1-p_2\), with Group 1 minus Group 2. Then examine both endpoints. If the lower endpoint is greater than zero, every value in the interval is positive, so the interval supports the conclusion that \(p_1>p_2\). If the upper endpoint is less than zero, every value is negative, so it supports the conclusion that \(p_1<p_2\), or equivalently that Group 2 has the higher proportion.
The interval’s direction depends on its parameter order. For example, an interval of \((0.03,0.25)\) for \(p_1-p_2\) favors Group 1. If the subtraction order were reversed, the corresponding interval for \(p_2-p_1\) would be \((-0.25,-0.03)\), which favors Group 2. The conclusion about which group is higher does not change; only the sign changes with the order.
- If \(0<L<U\), where \(L\) and \(U\) are the interval endpoints for \(p_1-p_2\), the interval is entirely positive and Group 1 has the higher population proportion according to the interval.
- If \(L<U<0\), the interval is entirely negative and Group 2 has the higher population proportion according to the interval.
- If zero is in the interval, including at an endpoint, the interval does not establish a direction in the same way.
“Higher proportion” describes direction, not necessarily practical importance. An interval entirely above zero might contain only small positive differences, or it might include differences large enough to matter in the setting. Read the endpoints to describe the size of the plausible differences as well as their direction.
Worked Examples
Worked Example: An Interval Entirely Above Zero
Setting: Imagine independent random samples of residents from two neighborhoods. A success is a resident who uses a public library at least once a month. In Neighborhood 1, 99 of 150 sampled residents report monthly use; in Neighborhood 2, 78 of 150 do. Assume each neighborhood has at least 1,500 residents. Construct and interpret a 95% confidence interval for \(p_1-p_2\), with Neighborhood 1 first.
State: Let \(p_1\) and \(p_2\) be the true proportions of residents in Neighborhoods 1 and 2, respectively, who use a public library at least once a month. The interval estimates \(p_1-p_2\).
Plan: The two samples are independent random samples, supporting the Random condition and independence between groups. Each sample of 150 is at most 10% of its population because each neighborhood has at least 1,500 residents. For the Large Counts condition, Neighborhood 1 has 99 successes and \(150-99=51\) failures; Neighborhood 2 has 78 successes and \(150-78=72\) failures. All four counts are at least 10.
Do: Calculate the two sample proportions and their difference:
Use the unpooled standard error for the confidence interval:
For 95% confidence, \(z^*\approx1.959964\). The margin of error is \(1.959964(0.0562139)\approx0.11018\). Therefore:
Conclude: We are 95% confident that the true proportion of residents who use a public library monthly in Neighborhood 1 minus the true proportion in Neighborhood 2 is between about 0.030 and 0.250. The entire interval is above zero, so it supports the conclusion that the population proportion is higher in Neighborhood 1. The interval describes plausible differences of about 3.0 to 25.0 percentage points; it does not mean there is a 95% probability that Neighborhood 1’s proportion is higher.
Worked Example: An Interval Entirely Below Zero
Setting: Imagine independent random samples of customers from two grocery stores. A success is a customer who brings a reusable shopping bag. Store 1 has 54 successes among 120 sampled customers, and Store 2 has 78 among 120. Assume each store serves at least 1,200 customers in the population of interest. Find and interpret a 95% confidence interval for \(p_1-p_2\), with Store 1 first.
State: Let \(p_1\) and \(p_2\) be the true proportions of customers at Stores 1 and 2, respectively, who bring a reusable shopping bag. The parameter is \(p_1-p_2\).
Plan: The samples are independent random samples, supporting randomness and independence between groups. Each sample of 120 is at most 10% of its store’s customer population because each population is at least 1,200. Store 1 has 54 successes and \(120-54=66\) failures; Store 2 has 78 successes and \(120-78=42\) failures. Each of the four counts is at least 10, so the Large Counts condition is met.
Do: The point estimate is:
Calculate the standard error and the 95% interval:
Conclude: We are 95% confident that the true proportion of customers bringing a reusable bag at Store 1 minus the true proportion at Store 2 is between about \(-0.323\) and \(-0.077\). The entire interval is below zero, so it supports the conclusion that Store 2 has the higher population proportion. In the reverse order, Store 2 minus Store 1, the interval would be approximately \((0.077,0.323)\); both versions describe the same direction.
Worked Example: Direction and Size in Percentage Points
Setting: Imagine independent random samples of 100 households from each of two towns. A success is a household that has a working smoke alarm on every floor. Town 1 has 64 successes and Town 2 has 48. Assume each town has at least 1,000 households. Find and interpret a 90% confidence interval for \(p_1-p_2\), with Town 1 first.
State and plan: Let \(p_1\) and \(p_2\) be the true proportions of households in Towns 1 and 2, respectively, with a working smoke alarm on every floor. The independent random samples support the Random condition and independence between groups. Each sample of 100 is at most 10% of its population. The success and failure counts are 64 and 36 in Town 1, and 48 and 52 in Town 2. All four counts are at least 10, so the interval’s Large Counts condition is met.
Do: The estimated difference is \(0.64-0.48=0.16\). The standard error is:
For 90% confidence, \(z^*\approx1.644854\). The margin of error is \(1.644854(0.0692820)\approx0.11396\), so:
Conclude: We are 90% confident that the true proportion of households with a working smoke alarm on every floor in Town 1 minus the true proportion in Town 2 is between about 0.046 and 0.274. The interval is entirely positive, supporting the conclusion that Town 1’s population proportion is higher. In context, the plausible differences range from about 4.6 to 27.4 percentage points. The positive direction does not, by itself, tell us whether every difference in that range is practically important.
Common Mistakes and AP Exam Tips
- Reading the sign without checking the order: A positive interval for \(p_1-p_2\) favors Group 1, but a positive interval for \(p_2-p_1\) favors Group 2. Name the groups in the parameter definition before interpreting the endpoints.
- Using only the point estimate: A positive sample difference does not by itself establish that \(p_1>p_2\). For this interpretation, examine the confidence interval: its lower endpoint must be above zero for it to be entirely positive.
- Saying “the interval proves” one group is higher: A confidence interval gives plausible values based on a method with a stated confidence level. A full-credit conclusion says the interval supports the direction that one population proportion is higher, rather than claiming certainty or proof.
- Confusing a proportion difference with a percent change: A difference such as 0.10 is 10 percentage points, not automatically a 10% increase. State the interval in the units that fit the context and do not call percentage points a percent change.
- Ignoring a rounded endpoint near zero: If an endpoint rounds to 0.000, use the unrounded endpoint to decide whether the interval is entirely positive, entirely negative, or includes zero. Do not make a directional claim based only on a rounded display that hides the sign.
- Claiming practical importance from direction alone: An interval entirely above zero identifies a direction, but the endpoints determine the plausible sizes. Discuss whether those sizes matter in context only when the question provides a practical benchmark or asks for that judgment.
Key Takeaway
For an interval estimating \(p_1-p_2\), an interval entirely above zero supports the conclusion that Group 1 has the higher population proportion. An interval entirely below zero supports the conclusion that Group 2 has the higher population proportion. The subtraction order determines the sign, and the endpoints describe the plausible sizes of the difference.
Check Your Understanding
For each interval, assume the parameter is \(p_1-p_2\). Use the sign and endpoints to explain the direction and what the interval does—and does not—establish.
- An interval is \((0.02,0.18)\). Which group has the higher population proportion according to the interval?
- An interval is \((-0.21,-0.04)\). Which group has the higher population proportion, and why?
- What direction would an interval for \(p_2-p_1\) have if the interval for \(p_1-p_2\) were \((0.05,0.12)\)?
- Why is it not enough to observe that \(\hat{p}_1-\hat{p}_2\) is positive when deciding whether the interval supports \(p_1>p_2\)?
- An interval’s lower endpoint rounds to 0.000. What should you check before deciding whether it is entirely positive?