Tutorials › AP Statistics › Confidence Intervals for Differences in Experiments

Two-proportion confidence intervals · Tutorial 513 of 1000

Confidence Intervals for Differences in Experiments

Use a two-proportion confidence interval to estimate the treatment-minus-control difference and explain what the interval says about an experiment’s effect.

Intermediate 9 min read

What You'll Learn

  • Define the treatment-minus-control difference in population proportions.
  • Check the conditions for a two-proportion confidence interval in a randomized experiment.
  • Calculate and interpret the interval in percentage points.
  • Explain when random assignment supports a causal interpretation.
  • Distinguish conclusions about experimental participants from generalizations to a wider population.

From a Difference in Proportions to an Experimental Effect

In Choosing Which Group Is Group 1, you learned to state the subtraction order and keep it consistent. For a randomized experiment comparing a treatment with a control, a natural choice is treatment minus control. A confidence interval in this order estimates how much higher or lower the treatment group’s success proportion is than the control group’s.

Random assignment matters to the interpretation. It uses chance to place experimental units in treatment and control groups, helping make the groups comparable before the treatment is applied. If the interval supports a difference, the design can support a cause-and-effect interpretation of the treatment for the experimental units studied. But random assignment alone does not make those units representative of a larger population.

Definition: Let \(p_T\) be the true proportion of experimental units with the defined success under the treatment condition, and let \(p_C\) be the true proportion with that success under the control condition. The treatment-minus-control difference is \(p_T-p_C\), estimated by \(\hat{p}_T-\hat{p}_C\). A positive value means the treatment proportion is higher; a negative value means the control proportion is higher.

The two-proportion \(z\)-interval procedure from Constructing a Two-Proportion z-Interval by Hand applies here. Its point estimate is the treatment sample proportion minus the control sample proportion. Its standard error uses the two observed proportions separately; do not pool them for a confidence interval. The interval gives plausible values for the difference in population proportions, in the stated order.

Formula: For a confidence level with critical value \(z^*\), calculate the interval as:
$$ (\hat{p}_T-\hat{p}_C)\mathbin{\pm}z^* \sqrt{\frac{\hat{p}_T(1-\hat{p}_T)}{n_T} +\frac{\hat{p}_C(1-\hat{p}_C)}{n_C}} $$
Here, \(n_T\) and \(n_C\) are the numbers of experimental units in the treatment and control groups. The interval estimates \(p_T-p_C\), not a relative percent change.

Conditions in a Randomized Experiment

As explained in Two-Sample Independence Conditions for Proportions and Large Counts for Each Group in Two-Proportion Intervals, check the study design, independence, and counts before calculating. An experiment’s random assignment supports the chance-based design condition. The treatment and control groups must be separate, and each experimental unit should contribute one outcome to only one group. If units were randomly sampled without replacement from a finite population, check the 10% condition using the total sample, not each assigned group separately.

Conditions:
  • Random: Experimental units are randomly assigned to treatment and control, or the study uses appropriate random samples. Random assignment supports a causal interpretation; random sampling supports generalization to the population sampled.
  • Independent groups and observations: Each unit is in only one group, and the outcomes can reasonably be treated as independent. If sampling without replacement, the total sample must be no more than 10% of the population.
  • Large Counts: Each group has at least 10 successes and at least 10 failures.

For a randomized experiment, it is important to say what the design does and does not justify. Random assignment can support a causal conclusion about the treatment, assuming the experiment was carried out as described. Generalizing that conclusion to people beyond the experimental units requires a basis for representativeness, such as random sampling from that wider population. Do not claim that random assignment by itself provides such a basis.

Once a valid interval is calculated, use its endpoints to discuss both direction and size. If the entire interval is above zero, plausible differences favor treatment. If it is entirely below zero, plausible differences favor control. If it contains zero, no difference remains among the plausible values; that does not prove the treatment and control proportions are equal. These sign interpretations build on What It Means When the Interval Contains Zero and Intervals Entirely Above or Below Zero.

Worked Examples

Worked Example: A Treatment Group Has a Higher Success Proportion

Setting: Imagine 300 students randomly selected from a school district with at least 3,000 students. They are randomly assigned in equal numbers to use either a study-planning app or a standard weekly planning sheet. A success is completing a specified set of assignments on time during the next month. In the app group, 108 of 150 students succeed; in the control group, 87 of 150 succeed. Construct and interpret a 95% confidence interval for treatment minus control.

State: Let \(p_T\) and \(p_C\) be the true proportions of students in the population of interest who would complete the specified assignments on time under the app and control conditions, respectively. The parameter is \(p_T-p_C\), with treatment as Group 1 and control as Group 2.

Plan: Students were randomly selected and randomly assigned, supporting the Random condition. The groups are separate, and each student contributes one outcome. The total sample of 300 is at most 10% of the district population because \(300\leq0.10(3000)=300\). For Large Counts, the app group has 108 successes and \(150-108=42\) failures; the control group has 87 successes and \(150-87=63\) failures. All four counts are at least 10.

Do: The sample proportions and difference are:

$$ \hat{p}_T-\hat{p}_C =\frac{108}{150}-\frac{87}{150} =0.72-0.58 =0.14 $$

The standard error and margin of error are:

$$ \begin{aligned} SE_{\hat{p}_T-\hat{p}_C} &=\sqrt{\frac{0.72(0.28)}{150}+\frac{0.58(0.42)}{150}}\\ &=\sqrt{0.001344+0.001624}\\ &=\sqrt{0.002968}\approx0.054479 \end{aligned} $$

For 95% confidence, \(z^*\approx1.959964\). The margin of error is \(1.959964(0.054479)\approx0.10678\). Therefore:

$$ 0.14\mathbin{\pm}0.10678 \quad\Longrightarrow\quad (0.03322,\ 0.24678) $$

Conclude: We are 95% confident that the true on-time completion proportion under the app condition minus the true proportion under the control condition is between about 0.033 and 0.247. The interval is entirely positive, so it supports the conclusion that the app increased the on-time completion proportion for students like those studied. The plausible increase is about 3.3 to 24.7 percentage points. Because students were randomly assigned, a causal interpretation for the experimental units is reasonable; random selection also supports generalizing to the district population, subject to the study’s implementation and other limitations.

Worked Example: An Interval That Includes Zero

Setting: Imagine 400 randomly selected household water systems from a region with at least 4,000 systems. In a randomized experiment, 200 receive a new filter and 200 receive the standard filter. A success is meeting a defined water-quality target after one month. The new-filter group has 124 successes, and the standard-filter group has 110. Construct and interpret a 90% confidence interval for new filter minus standard filter.

State: Let \(p_N\) and \(p_S\) be the true proportions of water systems meeting the target under the new and standard filters, respectively. The parameter is \(p_N-p_S\).

Plan: Random selection and random assignment support the Random condition. Each water system is assigned to one filter, so the groups are separate. The total sample of 400 is at most 10% of the region’s population because \(400\leq0.10(4000)=400\). For Large Counts, the new-filter group has 124 successes and \(200-124=76\) failures; the standard-filter group has 110 successes and \(200-110=90\) failures. Each count is at least 10.

Do: Calculate the estimated difference:

$$ \hat{p}_N-\hat{p}_S =\frac{124}{200}-\frac{110}{200} =0.62-0.55 =0.07 $$

The standard error is:

$$ \begin{aligned} SE_{\hat{p}_N-\hat{p}_S} &=\sqrt{\frac{0.62(0.38)}{200}+\frac{0.55(0.45)}{200}}\\ &=\sqrt{0.001178+0.0012375}\\ &=\sqrt{0.0024155}\approx0.049148 \end{aligned} $$

For 90% confidence, \(z^*\approx1.644854\). The margin of error is \(1.644854(0.049148)\approx0.08084\), so:

$$ 0.07\mathbin{\pm}0.08084 \quad\Longrightarrow\quad (-0.01084,\ 0.15084) $$

Conclude: We are 90% confident that the difference in the true proportions meeting the water-quality target, new filter minus standard filter, is between about \(-0.011\) and \(0.151\). Since the interval contains zero, the data are compatible with a slightly lower proportion under the new filter, no difference, or a higher proportion. The interval does not establish that the filters are equally effective. Random assignment permits a causal interpretation of the comparison for the experimental units, but the interval alone does not show that the new filter is better.

Worked Example: Report the Size of a Positive Effect

Setting: Imagine 500 randomly selected plants from a large greenhouse operation are assigned at random to either a new watering schedule or the usual schedule. A success is producing a marketable flower within a specified period. The new-schedule group has 175 successes among 250 plants; the usual-schedule group has 150 among 250. Assume the operation has at least 5,000 plants. Construct a 95% confidence interval for new schedule minus usual schedule.

State: Let \(p_N\) and \(p_U\) be the true proportions of plants producing a marketable flower under the new and usual schedules, respectively. The parameter is \(p_N-p_U\).

Plan: The plants are randomly selected and randomly assigned, supporting the Random condition. The groups are separate, and each plant is measured once. The total sample of 500 is at most 10% of the operation’s population because \(500\leq0.10(5000)=500\). The new-schedule group has 175 successes and \(250-175=75\) failures; the usual-schedule group has 150 successes and \(250-150=100\) failures. All counts are at least 10.

Do: The estimate of the difference is:

$$ \hat{p}_N-\hat{p}_U =\frac{175}{250}-\frac{150}{250} =0.70-0.60 =0.10 $$

The standard error is:

$$ \begin{aligned} SE_{\hat{p}_N-\hat{p}_U} &=\sqrt{\frac{0.70(0.30)}{250}+\frac{0.60(0.40)}{250}}\\ &=\sqrt{0.00084+0.00096}\\ &=\sqrt{0.0018}\approx0.042426 \end{aligned} $$

For 95% confidence, \(z^*\approx1.959964\), and the margin of error is \(1.959964(0.042426)\approx0.08315\). The interval is:

$$ 0.10\mathbin{\pm}0.08315 \quad\Longrightarrow\quad (0.01685,\ 0.18315) $$

Conclude: We are 95% confident that the true marketable-flower proportion under the new schedule minus the true proportion under the usual schedule is between about 0.017 and 0.183. This corresponds to a plausible increase of about 1.7 to 18.3 percentage points. Since the interval is above zero and plants were randomly assigned, the results support a positive causal effect of the new schedule for the experimental units. The interval does not imply that every plant benefits or that the increase is large enough to be practically important.

Common Mistakes and AP Exam Tips

  • Reversing the subtraction order: If the question asks for treatment minus control, define the parameter as \(p_T-p_C\) and calculate \(\hat{p}_T-\hat{p}_C\). A positive result then favors treatment.
  • Checking the 10% condition separately for each group: When units were sampled as one group before assignment, check the total sample. For example, 150 units in each arm means a total sample of 300, so verify that 300 is no more than 10% of the population.
  • Claiming random assignment allows generalization: Random assignment supports a causal interpretation; random sampling supports generalization. State which feature the experiment actually used.
  • Calling an interval containing zero proof of no effect: Zero is one plausible difference, but other values in the interval may represent meaningful increases or decreases. Say that the interval includes no difference, not that the proportions are equal.
  • Reporting a difference as a percent increase: A difference of 0.10 is 10 percentage points. It is not automatically a 10% increase relative to the control proportion.
  • Giving a conclusion without naming the outcome: A full-credit interpretation identifies the success, both conditions, the subtraction order, the plausible values, and the confidence level.
AP Exam Tip: In a randomized experiment, write the interval interpretation first as a statement about the treatment-minus-control difference in population proportions. Then use the interval’s position relative to zero to describe direction. Finally, connect random assignment to causation and distinguish that from any claim about generalizing to a wider population.

Key Takeaway

A two-proportion confidence interval in a randomized experiment estimates the treatment-minus-control difference in success proportions. Its sign and endpoints describe plausible effect sizes, while the study design determines whether a causal interpretation or broader generalization is justified.

Key takeaway: Define the treatment and control proportions, check the experiment’s conditions, and interpret the interval in percentage points. Random assignment can support a causal conclusion for the experimental units; generalizing beyond them requires an appropriate sampling basis.

Check Your Understanding

For each question, preserve the treatment-minus-control order and distinguish what the interval says from what the experimental design supports.

  1. An experiment has 90 successes among 120 treatment units and 72 among 120 control units. What is the sample difference in proportions, in percentage points?
  2. If the total sample of 240 units was randomly selected without replacement, what minimum population size would satisfy the 10% condition?
  3. A confidence interval for \(p_T-p_C\) is \((0.02,0.14)\). What does its sign indicate about the treatment group’s success proportion?
  4. A confidence interval for treatment minus control is \((-0.04,0.10)\). What values remain plausible, and why would it be incorrect to conclude that the conditions have equal proportions?
  5. What does random assignment support, and what additional design feature would support generalizing the result to a wider population?