Tutorials › AP Statistics › Confidence Intervals for Differences in Observational Studies

Two-proportion confidence intervals · Tutorial 514 of 1000

Confidence Intervals for Differences in Observational Studies

Use a two-proportion confidence interval to describe plausible differences between observational groups while keeping conclusions about association separate from claims about cause and effect.

Intermediate 10 min read

What You'll Learn

  • Define the population difference estimated by an interval from two observational groups
  • Interpret endpoints and signs using the stated subtraction order
  • Check random sampling, independence, the 10% condition, and Large Counts
  • Distinguish generalizing to sampled populations from making causal claims
  • Explain why a confidence interval cannot remove confounding or selection bias

What a Difference in an Observational Study Estimates

In Confidence Intervals for Differences in Experiments, random assignment was central to interpreting a difference as a possible treatment effect. An observational study is different: researchers observe people or other units in groups that already exist, rather than assigning them to groups by chance. A confidence interval can estimate a difference between those groups, but the interval by itself cannot show that group membership caused the difference.

For example, researchers might compare the proportion of adults who meet a sleep target among those who usually drink caffeine in the evening and those who do not. The observed proportions can differ, and a confidence interval can describe plausible values for the difference between the corresponding population proportions. But other characteristics—such as work schedules or health habits—could be related to both evening caffeine use and sleep.

Definition: Let \(p_1\) and \(p_2\) be the true proportions with the same defined outcome in two populations or naturally occurring groups. The difference \(p_1-p_2\) is estimated by \(\hat{p}_1-\hat{p}_2\). In an observational study, this parameter describes a difference between groups; it is not automatically a treatment effect or a causal effect.

The two-proportion \(z\)-interval from Constructing a Two-Proportion z-Interval by Hand estimates \(p_1-p_2\). Use the group order named in the question, or state your chosen order clearly. A positive difference means the proportion is higher in Group 1; a negative difference means it is higher in Group 2. The units of the difference are proportions, often reported as percentage points.

Formula: For a confidence level with critical value \(z^*\), the interval is:
$$ (\hat{p}_1-\hat{p}_2)\mathbin{\pm}z^* \sqrt{\frac{\hat{p}_1(1-\hat{p}_1)}{n_1} +\frac{\hat{p}_2(1-\hat{p}_2)}{n_2}} $$
The standard error uses each group’s sample proportion separately. As in earlier tutorials, do not pool the sample proportions when constructing a confidence interval.

Conditions and the Scope of the Conclusion

An observational study can support inference about population proportions when its data collection and sample sizes support the interval procedure. As in Two-Sample Independence Conditions for Proportions and Large Counts for Each Group in Two-Proportion Intervals, check the design and counts before calculating. For two independent random samples, the Random condition comes from the samples being randomly selected from their respective populations—not from random assignment.

Conditions:
  • Random: Each group should be represented by an appropriate random sample from its target population. If the samples are not random, generalizing to those populations may not be justified.
  • Independent groups and observations: The samples should be separate, and each sampled unit should contribute one outcome to only one group. If a sample is drawn without replacement, check the 10% condition separately for each sample: each sample size must be no more than 10% of its population.
  • Large Counts: In each group, there must be at least 10 observed successes and at least 10 observed failures.

When these conditions are met, the interval estimates the difference between the population proportions for the two groups represented by the samples. Random sampling can support generalizing to those populations. It does not make the groups comparable in every other respect, however, and it does not rule out confounding variables. In an observational study, a factor associated with both group membership and the outcome may help explain an observed difference.

This is the key contrast with an experiment. Random assignment can help balance other factors across treatment groups and can support a causal interpretation. In an observational study, group membership was not assigned at random, so even an interval entirely above or below zero is evidence of a difference between the group proportions—not, by itself, evidence that one group’s behavior caused the outcome.

Interpret the interval itself carefully, too. A 95% confidence interval gives a range of plausible values for the specified population difference, in the stated order. If it contains zero, no difference is among the plausible values; that does not prove the proportions are equal. If the whole interval is positive or negative, it indicates a direction for the difference between the population proportions. None of these conclusions changes the observational nature of the study.

Worked Examples

Worked Example: Evening Caffeine Use and Meeting a Sleep Target

Setting: Imagine researchers independently select random samples of adults who usually drink caffeine after 6 p.m. and adults who do not. Each sampled adult reports whether they usually sleep at least seven hours on work nights. Among 200 evening-caffeine users, 120 meet the target; among 200 nonusers, 96 meet it. Each group’s population contains at least 2,000 adults. Construct and interpret a 95% confidence interval for evening-caffeine users minus nonusers.

State: Let \(p_1\) be the true proportion of adults who usually drink caffeine after 6 p.m. and sleep at least seven hours on work nights. Let \(p_2\) be the corresponding proportion among adults who do not usually drink caffeine after 6 p.m. The parameter is \(p_1-p_2\).

Plan: The two groups were independently randomly sampled from their respective populations, supporting the Random condition. The groups are separate, and each adult contributes one response. Each sample size is at most 10% of its population because \(200\leq0.10(2000)=200\). For Large Counts, the evening-caffeine group has 120 successes and \(200-120=80\) failures; the nonuser group has 96 successes and \(200-96=104\) failures. All four counts are at least 10.

Do: Calculate the sample proportions and their difference:

$$ \hat{p}_1-\hat{p}_2 =\frac{120}{200}-\frac{96}{200} =0.60-0.48 =0.12 $$

The standard error is:

$$ \begin{aligned} SE_{\hat{p}_1-\hat{p}_2} &=\sqrt{\frac{0.60(0.40)}{200}+\frac{0.48(0.52)}{200}}\\ &=\sqrt{0.001200+0.001248}\\ &=\sqrt{0.002448}\approx0.049477 \end{aligned} $$

For 95% confidence, \(z^*\approx1.959964\). The margin of error is \(1.959964(0.049477)\approx0.09697\). Thus:

$$ 0.12\mathbin{\pm}0.09697 \quad\Longrightarrow\quad (0.02303,\ 0.21697) $$

Conclude: We are 95% confident that the true proportion of adults who usually meet the sleep target among evening-caffeine users minus the true proportion among nonusers is between about 0.023 and 0.217. In this group order, the interval suggests a difference of about 2.3 to 21.7 percentage points, with the proportion higher among evening-caffeine users. This is an observational comparison: it does not show that evening caffeine improves sleep. Other differences between the groups could be related to both caffeine habits and sleep.

Worked Example: A Difference That Could Go Either Way

Setting: Imagine researchers take separate random samples of adults who usually work an evening shift and adults who usually work a daytime shift. The outcome is usually getting at least seven hours of sleep before a workday. In the evening-shift sample, 84 of 160 adults meet the target; in the daytime-shift sample, 72 of 160 do. Each target population contains at least 1,600 adults. Construct and interpret a 90% confidence interval for evening-shift workers minus daytime-shift workers.

State: Let \(p_E\) and \(p_D\) be the true proportions who meet the sleep target among evening-shift and daytime-shift workers, respectively. The parameter is \(p_E-p_D\).

Plan: Both groups were randomly sampled from their respective populations, supporting the Random condition. They are separate groups, with one response per adult. The 10% condition is met for each sample because \(160\leq0.10(1600)=160\). For Large Counts, the evening-shift sample has 84 successes and \(160-84=76\) failures; the daytime-shift sample has 72 successes and \(160-72=88\) failures. Every count is at least 10.

Do: The sample proportions and difference are:

$$ \hat{p}_E-\hat{p}_D =\frac{84}{160}-\frac{72}{160} =0.525-0.450 =0.075 $$

The standard error is:

$$ \begin{aligned} SE_{\hat{p}_E-\hat{p}_D} &=\sqrt{\frac{0.525(0.475)}{160}+\frac{0.450(0.550)}{160}}\\ &=\sqrt{0.0015586+0.0015469}\\ &=\sqrt{0.0031055}\approx0.055727 \end{aligned} $$

For 90% confidence, \(z^*\approx1.644854\). The margin of error is \(1.644854(0.055727)\approx0.09166\), so the interval is:

$$ 0.075\mathbin{\pm}0.09166 \quad\Longrightarrow\quad (-0.01666,\ 0.16666) $$

Conclude: We are 90% confident that the difference in the true proportions meeting the sleep target, evening-shift workers minus daytime-shift workers, is between about \(-0.017\) and \(0.167\). The interval includes zero, so the plausible differences include a slightly lower proportion among evening-shift workers, no difference, and a higher proportion. It does not establish that the proportions are equal. Since shift schedules were observed rather than randomly assigned, the interval also cannot show that a shift schedule caused a difference in sleep.

Worked Example: Transit Passes and Short Commutes

Setting: Imagine researchers independently select random samples of residents who have a public-transit pass and residents who do not. The outcome is having a commute of less than 30 minutes. In the pass-holder sample, 72 of 180 residents have a short commute; in the sample without passes, 90 of 200 do. The pass-holder population has at least 1,800 residents, and the population without passes has at least 2,000. Construct and interpret a 95% confidence interval for pass holders minus residents without passes.

State: Let \(p_P\) be the true proportion of public-transit pass holders with a commute under 30 minutes, and let \(p_N\) be the corresponding proportion among residents without a pass. The parameter is \(p_P-p_N\).

Plan: Separate random samples were drawn from each group’s population, supporting the Random condition. The samples are independent, and each resident contributes one outcome. The 10% condition holds for both samples: \(180\leq0.10(1800)=180\) and \(200\leq0.10(2000)=200\). For Large Counts, pass holders have 72 successes and \(180-72=108\) failures; residents without passes have 90 successes and \(200-90=110\) failures. All counts are at least 10.

Do: The sample proportions and difference are:

$$ \hat{p}_P-\hat{p}_N =\frac{72}{180}-\frac{90}{200} =0.40-0.45 =-0.05 $$

The standard error is:

$$ \begin{aligned} SE_{\hat{p}_P-\hat{p}_N} &=\sqrt{\frac{0.40(0.60)}{180}+\frac{0.45(0.55)}{200}}\\ &=\sqrt{0.0013333+0.0012375}\\ &=\sqrt{0.0025708}\approx0.050703 \end{aligned} $$

For 95% confidence, \(z^*\approx1.959964\). The margin of error is \(1.959964(0.050704)\approx0.09938\). Therefore:

$$ -0.05\mathbin{\pm}0.09938 \quad\Longrightarrow\quad (-0.14938,\ 0.04938) $$

Conclude: We are 95% confident that the true proportion with a commute under 30 minutes among pass holders minus the true proportion among residents without passes is between about \(-0.149\) and \(0.049\). The interval includes zero and allows differences in either direction. Even if the interval had been entirely below zero, that would describe an association between pass-holder status and commute length, not prove that having a pass causes a longer commute. Neighborhood, workplace location, and other factors could be related to both pass ownership and commute time.

Common Mistakes and AP Exam Tips

  • Calling the difference a treatment effect: In an observational study, the groups were not formed by random assignment. Describe the difference in population proportions rather than claiming one group’s characteristic caused the outcome.
  • Confusing random sampling with random assignment: Random sampling can support generalizing to the sampled population. Random assignment can support a causal interpretation. They serve different purposes.
  • Ignoring the target populations: State which two populations the random samples represent. If the sampling method does not support representing those populations, a confidence interval does not fix that limitation.
  • Reversing the subtraction order: Define \(p_1-p_2\) before calculating. A negative interval or estimate means Group 2 has the higher proportion, not that the result is invalid.
  • Claiming equality when an interval contains zero: Zero is one plausible value, but nonzero values may also be plausible. Say that the interval includes no difference; do not say the study proves equal proportions.
  • Using causal wording for an association: “The proportion was higher among Group 1” describes a comparison. “Being in Group 1 increased the proportion” makes a causal claim that an observational interval alone does not justify.
AP Exam Tip: A complete interpretation names both groups, the outcome, the subtraction order, the confidence level, and the interval’s endpoints. Then state what the design supports: random sampling may justify generalizing to the target populations, but without random assignment the interval alone does not establish causation.

Key Takeaway

A two-proportion confidence interval from independent random samples estimates a difference between two population proportions. Its endpoints and sign describe plausible differences between the groups, while the study design limits what can be concluded about why those groups differ.

Key takeaway: Interpret the interval as an association between the specified populations and preserve the group order. Random sampling can support generalization; an observational study does not use random assignment to establish cause and effect.

Check Your Understanding

For each question, identify the population difference and distinguish an interval’s conclusion from what the observational design can support.

  1. In a random sample of 150 residents with a bike-share membership, 90 report using a helmet on their last ride. In an independent random sample of 150 residents without a membership, 75 do. What is the sample difference for membership holders minus nonholders?
  2. A 95% confidence interval for \(p_1-p_2\) is \((0.04,0.18)\). What does its sign say about the two population proportions in the stated order?
  3. A 90% confidence interval for \(p_1-p_2\) is \((-0.06,0.11)\). What differences remain plausible, and why is it incorrect to conclude that the population proportions are equal?
  4. Why does randomly sampling adults who already use a service and adults who do not use it fail to establish that using the service caused an outcome difference?
  5. If two independent samples are taken without replacement, what population-size check is needed for the 10% condition?