Tutorials › AP Statistics › Interpreting a Confidence Interval for p1 – p2

Two-proportion confidence intervals · Tutorial 509 of 1000

Interpreting a Confidence Interval for p1 - p2

Practice explaining what the endpoints of a confidence interval say about the difference between two population proportions and which group has the higher proportion.

Intermediate 9 min read

What You'll Learn

  • State what \(p_1-p_2\) represents in the situation.
  • Describe the interval’s endpoints as plausible values for the population difference.
  • Use the sign and group order to identify which population proportion is higher.
  • Convert a difference in proportions to percentage points when useful.
  • Explain confidence without claiming the parameter has a probability of being in the calculated interval.
  • Keep an interval interpretation distinct from a claim about individual outcomes.

From Calculator Endpoints to a Sentence

In Using 2-PropZInt on the Calculator, you used the TI-84 to find an interval estimating \(p_1-p_2\). The calculator supplies endpoints, but it cannot explain what those endpoints mean in the situation. A complete interpretation identifies the populations and characteristic, preserves the subtraction order, and describes the interval as plausible values for the difference in the two true proportions.

Recall from Comparing Two Proportions: Setting and Notation that \(p_1\) and \(p_2\) are the true proportions for the two defined groups. The interval estimates \(p_1-p_2\), not either proportion by itself. Its sign and size therefore describe how the first group’s proportion compares with the second group’s proportion.

Definition: An interpretation of a confidence interval for \(p_1-p_2\) describes, in context, the plausible values of the difference between the two population proportions, with Group 1’s proportion minus Group 2’s proportion.

A useful sentence pattern is: “We are [confidence level] confident that the true proportion of [success] in [Group 1] minus the true proportion of [success] in [Group 2] is between [lower endpoint] and [upper endpoint].” The endpoints are differences in proportions. For easier reading, you can express them as percentage points.

Key interpretation rule: Keep the subtraction order \(p_1-p_2\). If the interval endpoints are positive, the plausible differences put Group 1’s proportion above Group 2’s. If both endpoints are negative, they put Group 1’s proportion below Group 2’s. State the comparison in the setting rather than leaving the direction implicit.

What the Confidence Level Says

A 95% confidence level describes the long-run success rate of the interval method: if many pairs of samples were collected in the same way and a 95% interval were calculated from each pair, about 95% of those intervals would capture the true difference \(p_1-p_2\). In an interpretation of one calculated interval, the standard AP wording is “We are 95% confident that the true difference is between … and ….”

Do not say that there is a 95% probability that the fixed population difference lies in this particular interval. The parameter \(p_1-p_2\) is fixed; the interval varies from sample to sample. The confidence level describes the method’s long-run performance, not a probability assigned to the fixed parameter after the interval has been calculated.

When translating endpoints, remember that a proportion difference of \(0.08\) is 8 percentage points. It is not necessarily an 8% relative increase. For example, if one group’s proportion is 0.50 and the other’s is 0.42, their difference is 0.08, or 8 percentage points. For the interval interpretation, focus on the difference in proportions, not an unsupported claim about relative percent change.

A Reliable Interpretation Process

Before writing the sentence, check the parameter definition and group order from the question. Then look at the signs of both endpoints and translate the difference into plain language. As covered in Constructing a Two-Proportion z-Interval by Hand and Two-Sample Independence Conditions for Proportions, the interval is appropriate only when its conditions are supported. A polished interpretation does not replace those checks in a complete response.

1
Name the difference.
Identify the success characteristic and say that the interval estimates \(p_1-p_2\), with Group 1 first.
2
Read the endpoints as differences.
Do not describe them as the proportions in either group. They are lower and upper plausible values for the difference between the population proportions.
3
Translate direction and units.
Use the signs to say which group has the higher proportion across the interval’s plausible values. Give the endpoints as proportions or percentage points.
4
State confidence carefully.
Use “We are [confidence level] confident,” and identify the populations and characteristic in context.

Worked Examples

Worked Example: A Positive Difference in Library Use

Setting: Imagine independent random samples of residents from two districts. A success is a resident who visited a public library in the past month. In District 1, 92 of 150 sampled residents report a visit; in District 2, 66 of 150 do. Assume each district has at least 1,500 residents. Find and interpret a 95% confidence interval for \(p_1-p_2\), with District 1 first.

State: Let \(p_1\) and \(p_2\) be the true proportions of residents in Districts 1 and 2, respectively, who visited a public library in the past month. The parameter of interest is \(p_1-p_2\).

Plan: The data are stated to come from independent random samples, supporting the Random condition and independence between groups. Each sample of 150 is at most 10% of its population because each district has at least 1,500 residents. District 1 has 92 successes and \(150-92=58\) failures; District 2 has 66 successes and \(150-66=84\) failures. All four counts are at least 10, so the Large Counts condition is met.

Do: The sample proportions and their difference are:

$$ \hat{p}_1=\frac{92}{150}=0.61333,\qquad \hat{p}_2=\frac{66}{150}=0.44,\qquad \hat{p}_1-\hat{p}_2=0.17333 $$

For a two-proportion z-interval, the standard error uses each group’s sample proportion:

$$ \begin{aligned} SE_{\hat{p}_1-\hat{p}_2} &=\sqrt{\frac{(92/150)(58/150)}{150} +\frac{(66/150)(84/150)}{150}}\\ &=\sqrt{0.00158104+0.00164267}\\ &=\sqrt{0.00322370}\approx0.05678 \end{aligned} $$

For 95% confidence, \(z^*\approx1.96\). The margin of error is \(1.959964(0.05678)\approx0.1113\), so the interval is:

$$ 0.17333\mathbin{\pm}0.11128 \quad\Longrightarrow\quad (0.06205,\ 0.28462) $$

Conclude: We are 95% confident that the true proportion of residents who visited a public library in the past month in District 1 minus the true proportion in District 2 is between about 0.062 and 0.285. Equivalently, the plausible difference is about 6.2 to 28.5 percentage points, with the District 1 proportion higher. This comparison is about the population proportions, not a claim that every District 1 resident is more likely to visit than every District 2 resident.

Worked Example: A Negative Difference in Garden Practices

Setting: Imagine independent random samples of community gardeners from two regions. A success is a gardener who uses drip irrigation. In Region 1, 30 of 120 sampled gardeners use it; in Region 2, 84 of 140 do. Assume each region has at least ten times its sample size in gardeners. Find and interpret a 95% confidence interval for \(p_1-p_2\), with Region 1 first.

State: Let \(p_1\) and \(p_2\) be the true proportions of gardeners in Regions 1 and 2, respectively, who use drip irrigation. The interval estimates \(p_1-p_2\).

Plan: The samples are stated to be independent random samples, supporting randomness and independence between groups. Each sample is no more than 10% of its region’s population, by the population-size assumption. Region 1 has 30 successes and 90 failures; Region 2 has 84 successes and 56 failures. All four counts are at least 10, so the Large Counts condition is met.

Do: Calculate the sample proportions and their difference in the stated group order:

$$ \hat{p}_1=\frac{30}{120}=0.25,\qquad \hat{p}_2=\frac{84}{140}=0.60,\qquad \hat{p}_1-\hat{p}_2=-0.35 $$

The standard error is:

$$ \begin{aligned} SE_{\hat{p}_1-\hat{p}_2} &=\sqrt{\frac{0.25(0.75)}{120}+\frac{0.60(0.40)}{140}}\\ &=\sqrt{0.00156250+0.00171429}\\ &=\sqrt{0.00327679}\approx0.05724 \end{aligned} $$

Using \(z^*\approx1.96\), the margin of error is \(1.959964(0.05724)\approx0.1122\). Therefore:

$$ -0.35\mathbin{\pm}0.1122 \quad\Longrightarrow\quad (-0.4622,\ -0.2378) $$

Conclude: We are 95% confident that the true proportion of gardeners using drip irrigation in Region 1 minus the true proportion in Region 2 is between about \(-0.462\) and \(-0.238\). In context, the Region 1 proportion is plausibly about 23.8 to 46.2 percentage points lower than the Region 2 proportion. The negative endpoints make the direction clear because the difference was defined as Region 1 minus Region 2.

Worked Example: Reporting a 90% Interval in Percentage Points

Setting: Imagine independent random samples of patients from two clinics. A success is a patient who schedules a follow-up appointment before leaving. In Clinic 1, 110 of 160 sampled patients schedule one; in Clinic 2, 75 of 150 do. Assume each clinic serves at least ten times its sample size. Find and interpret a 90% confidence interval for \(p_1-p_2\), with Clinic 1 first.

State: Let \(p_1\) and \(p_2\) be the true proportions of patients at Clinics 1 and 2, respectively, who schedule a follow-up appointment before leaving. We want to estimate \(p_1-p_2\).

Plan: The data come from independent random samples, supporting the Random condition and independence between groups. Each sample is no more than 10% of its clinic’s population under the stated population-size assumption. Clinic 1 has 110 successes and 50 failures; Clinic 2 has 75 successes and 75 failures. Each count is at least 10, satisfying the Large Counts condition.

Do: The sample proportions and their difference are:

$$ \hat{p}_1=\frac{110}{160}=0.6875,\qquad \hat{p}_2=\frac{75}{150}=0.50,\qquad \hat{p}_1-\hat{p}_2=0.1875 $$

Calculate the standard error using the unrounded sample proportions:

$$ \begin{aligned} SE_{\hat{p}_1-\hat{p}_2} &=\sqrt{\frac{0.6875(0.3125)}{160}+\frac{0.50(0.50)}{150}}\\ &=\sqrt{0.00134277+0.00166667}\\ &=\sqrt{0.00300944}\approx0.05486 \end{aligned} $$

For 90% confidence, \(z^*\approx1.644854\). The margin of error is \(1.644854(0.05486)\approx0.09023\). Thus the interval is approximately:

$$ 0.1875\mathbin{\pm}0.09023 \quad\Longrightarrow\quad (0.0973,\ 0.2777) $$

Conclude: We are 90% confident that the true proportion of patients who schedule a follow-up appointment at Clinic 1 minus the true proportion at Clinic 2 is between about 0.097 and 0.278. In percentage-point terms, the Clinic 1 proportion is plausibly about 9.7 to 27.8 points higher. The confidence level belongs to the interval method; it does not mean there is a 90% probability that the fixed difference is inside this particular interval.

Common Mistakes and AP Exam Tips

  • Reversing the group order: An interval for \(p_1-p_2\) must be described as Group 1 minus Group 2. If the group order is reversed, the difference changes sign. A full-credit response names both groups and preserves the order.
  • Calling the endpoints the two group proportions: Endpoints such as 0.06 and 0.28 describe possible differences, not a range for \(p_1\) or for \(p_2\). Say “the difference between the true proportions” or state the subtraction explicitly.
  • Omitting the characteristic or populations: “The difference is between 0.06 and 0.28” is not a complete interpretation. Identify what counts as a success and which two populations the proportions describe.
  • Ignoring a negative sign: A negative interval for \(p_1-p_2\) means the first proportion is lower than the second. Translate the sign into a contextual comparison rather than saying only “the difference is negative.”
  • Mixing up percent and percentage points: A difference of 0.10 is 10 percentage points. Do not call it a 10% increase unless a relative percentage change has actually been calculated and is appropriate to the question.
  • Misstating confidence: Do not claim there is a 95% probability that the fixed difference lies in the observed interval. Use the standard “We are 95% confident” wording and connect the confidence level to the interval method.
  • Making claims about individuals: A confidence interval for a difference in population proportions compares groups overall. It does not say that each individual in one group has a particular outcome or that one individual’s outcome is caused by group membership.
AP Exam Tip: A full-credit interpretation identifies the success, both populations, the subtraction order, the interval endpoints, and the confidence level. Then use the sign to express direction in context. If you report percentage points, make clear that they refer to a difference between proportions.

Key Takeaway

An interval for \(p_1-p_2\) describes plausible values of a population-proportion difference in the order Group 1 minus Group 2. Positive values indicate a higher proportion in Group 1; negative values indicate a lower proportion in Group 1. Interpret the endpoints in context, distinguish percentage points from percentages, and use confidence language that describes the interval method rather than assigning probability to a fixed parameter.

Key takeaway: Say what the difference measures, preserve the group order, translate its sign into direction, and report the endpoints as plausible values for the population difference.

Check Your Understanding

For each prompt, focus on the context, direction, and meaning of the interval for \(p_1-p_2\).

  1. An interval for the difference in the proportions of two schools’ students who bike to school is \((0.03, 0.12)\), with School A first. Write an interpretation in percentage points.
  2. An interval for \(p_1-p_2\) is \((-0.21,-0.08)\). Which group’s population proportion is higher across the plausible values, and why?
  3. Why is “There is a 95% probability that the true difference is in this interval” not the preferred interpretation of a 95% confidence interval?
  4. What does a difference of \(0.07\) mean in percentage points? How is that different from saying “a 7% increase”?
  5. For an interval estimating the difference in two population proportions, what information should a complete contextual interpretation include?