Tutorials › AP Statistics › Two Independent Samples Versus Paired Data

Two-proportion confidence intervals · Tutorial 517 of 1000

Two Independent Samples Versus Paired Data

Learn how to recognize independent samples, identify paired data, and decide whether a two-proportion confidence interval fits the study design.

Intermediate 9 min read

What You'll Learn

  • Identify when two groups consist of separate, independent observational units.
  • Distinguish independent samples from repeated measurements on the same people.
  • Recognize matched pairs and explain why their outcomes are dependent.
  • Check how random sampling, random assignment, and group membership affect the design.
  • Decide when a two-proportion confidence interval is appropriate and when it is not.

Start With the Study Design

Before calculating a two-proportion confidence interval, ask how the observations were collected and how the two groups were formed. In Comparing Two Proportions: Setting and Notation, the parameter \(p_1-p_2\) was defined as a difference between two population proportions. In Two-Sample Independence Conditions for Proportions, the design conditions for the interval were introduced. This tutorial focuses on a crucial design question: are the two sets of observations independent, or are they paired?

Two groups can have different names and still fail to be independent. For example, “before” and “after” are different time points, but if the same people are measured at both times, each after-measurement is linked to that person’s before-measurement. On the other hand, two groups can be compared with a two-proportion interval when each individual contributes data to just one group and the groups were collected or formed in a way that supports independence.

Definition: Two groups are independent for a two-proportion comparison when observations in one group are not paired or otherwise linked to particular observations in the other group. Each individual contributes to only one group, and the study design supports treating the group samples as independent.

The formula for a two-proportion interval estimates \(p_1-p_2\) using the uncertainty from each sample. Its standard error combines two group-specific variance estimates. That calculation relies on the groups being independent; it is not a general formula for any two sets of proportions.

$$ SE_{\hat{p}_1-\hat{p}_2} = \sqrt{ \frac{\hat{p}_1(1-\hat{p}_1)}{n_1} + \frac{\hat{p}_2(1-\hat{p}_2)}{n_2} } $$

If the same people provide both outcomes, or if people in one group are deliberately matched with people in the other, the observations are linked. The formula above does not account for that link. The issue is the design, not whether the sample sizes are equal or whether the two sample proportions look reasonable.

A Practical Design Check

Use the following questions before deciding that a two-proportion interval applies:

1
Identify the observational units.
These might be people, households, schools, or other entities. Be specific about what one observation represents.
2
Ask whether each unit appears in one group or both.
If the same unit supplies both outcomes, the data are paired rather than two independent samples.
3
Look for deliberate matching.
Pairs may be formed by matching people with similar characteristics, such as age or starting skill. A matched pair’s two outcomes are linked even though they come from different people.
4
Check how groups were obtained or formed.
Separate random samples can support a comparison of population proportions. In an experiment, random assignment to separate groups supports comparing the treatment and control proportions.
5
Only then check the interval conditions.
For an independent-groups interval, also check the appropriate randomization or sampling conditions, the 10% condition when sampling without replacement, and Large Counts for both groups, as covered in earlier tutorials.

A “yes” to separate groups is not enough by itself. For instance, two convenience groups may be separate but still lack a chance-based design that supports generalizing to populations. Conversely, a study may have random assignment but still use a paired design if each participant receives both treatments and provides both outcomes. Design labels are clues; the actual way each observation was produced is what matters.

Worked Examples

Worked Example: Two Independent Random Samples

Setting: A fictional environmental survey compares whether residents in two separate towns have installed a rain barrel. A random sample of 120 residents is selected from Town A, and a separate random sample of 100 residents is selected from Town B. In Town A, 72 sampled residents have a rain barrel; in Town B, 48 do. Is a two-proportion interval appropriate, and what is the 95% interval for \(p_1-p_2\), Town A minus Town B?

State: Let \(p_1\) be the true proportion of Town A residents with a rain barrel, and let \(p_2\) be the true proportion of Town B residents with a rain barrel. The parameter of interest is \(p_1-p_2\).

Plan: The towns were sampled separately, and each resident appears in only one sample, so the groups are independent. The samples are random. Assume Town A has at least 1,200 residents and Town B at least 1,000, so each sample is at most 10% of its town: \(120\leq0.10(1200)\) and \(100\leq0.10(1000)\). For Large Counts, Town A has 72 successes and \(120-72=48\) failures; Town B has 48 successes and \(100-48=52\) failures. All four counts are at least 10. The conditions support using a two-proportion \(z\)-interval.

Do: The sample proportions are \(\hat{p}_1=72/120=0.60\) and \(\hat{p}_2=48/100=0.48\). The point estimate is \(0.60-0.48=0.12\). Using \(z^*=1.959964\) for 95% confidence, calculate the standard error:

$$ \begin{aligned} SE_{\hat{p}_1-\hat{p}_2} &=\sqrt{\frac{0.60(0.40)}{120}+\frac{0.48(0.52)}{100}}\\ &=\sqrt{0.002000+0.002496}\\ &=\sqrt{0.004496}\approx0.067052 \end{aligned} $$

The margin of error is \(1.959964(0.067052)\approx0.131420\). Therefore, the interval is:

$$ 0.12\mathbin{\pm}0.131420 = (-0.0114,\ 0.2514) $$

Conclude: We are 95% confident that the true difference in the proportions of residents with a rain barrel, Town A minus Town B, is between about \(-0.0114\) and \(0.2514\). Because the samples are independent and the other conditions are met, the two-proportion interval is appropriate. The interval includes zero, so no difference remains among its plausible values.

Worked Example: The Same People Measured Twice

Setting: A fictional project asks 80 residents the same recycling question before and after a short information session. The outcome is whether each resident answers correctly. The paired results are: 30 answer correctly both times, 18 change from correct before to incorrect after, 26 change from incorrect before to correct after, and 6 answer incorrectly both times. Is a two-proportion interval for before minus after appropriate?

State: The question is whether the before-and-after data can be treated as two independent samples of proportions. The same 80 residents provide both responses, so each person’s responses are linked.

Plan: The two sets of measurements are paired, not independent. Each resident’s after-response is connected to that resident’s before-response. Therefore, the independence requirement for the two-proportion interval is not met. The other interval conditions cannot make this paired design into two independent samples.

Do: First verify that the transition counts represent the stated 80 residents: \(30+18+26+6=80\). The number correct before the session is the 30 correct both times plus the 18 who change from correct to incorrect, for \(30+18=48\). The number correct after the session is the 30 correct both times plus the 26 who change from incorrect to correct, for \(30+26=56\). Thus, the marginal sample proportions are \(48/80=0.60\) before and \(56/80=0.70\) after, a sample difference of \(0.10\) for after minus before.

Those marginal proportions do not erase the pairing. The data record which responses came from the same person, and the two-proportion standard-error formula assumes independent groups rather than using those person-level links. A method designed for paired binary outcomes would be needed to make an inference from these data; the independent two-proportion interval is not that method.

Conclude: Do not use a two-proportion confidence interval for these before-and-after measurements. The 80 residents form one group measured twice, not two independent samples. The sample proportions can be described, but the independent-groups interval formula does not apply to this design.

Worked Example: Separate Groups in a Randomized Experiment

Setting: In a fictional experiment, 200 volunteers are randomly assigned to one of two reminder systems for a study-planning app. Each volunteer uses only the assigned system. Of the 100 assigned to System 1, 90 submit a weekly plan; of the 100 assigned to System 2, 72 do. Is a two-proportion interval appropriate for estimating the treatment-minus-control difference in success proportions?

State: Let \(p_1\) be the true proportion of comparable experimental units who would submit a weekly plan using System 1, and \(p_2\) the corresponding proportion using System 2. The parameter is \(p_1-p_2\), System 1 minus System 2.

Plan: Volunteers were randomly assigned, and each volunteer used only one system, so the treatment groups are separate rather than paired. Random assignment supports the chance-based comparison; assume no volunteer’s outcome affects another volunteer’s outcome. Because the data come from random assignment rather than sampling without replacement from a finite population, a 10% sampling condition is not needed here. For Large Counts, System 1 has 90 successes and \(100-90=10\) failures; System 2 has 72 successes and \(100-72=28\) failures. All four counts are at least 10, so the Large Counts condition holds.

Do: The sample proportions are \(\hat{p}_1=90/100=0.90\) and \(\hat{p}_2=72/100=0.72\), giving a point estimate of \(0.18\). For a 95% interval, use \(z^*=1.959964\):

$$ \begin{aligned} SE_{\hat{p}_1-\hat{p}_2} &=\sqrt{\frac{0.90(0.10)}{100}+\frac{0.72(0.28)}{100}}\\ &=\sqrt{0.000900+0.002016}\\ &=\sqrt{0.002916}=0.054 \end{aligned} $$

The margin of error is \(1.959964(0.054)\approx0.105838\), so the interval is \(0.18\mathbin{\pm}0.105838=(0.0742,\ 0.2858)\), rounded to four decimal places.

Conclude: A two-proportion interval is appropriate because each volunteer contributed an outcome to only one randomly assigned group and the Large Counts condition is met. We are 95% confident that the true difference in weekly-plan submission proportions, System 1 minus System 2, is between about \(0.0742\) and \(0.2858\) for the experimental setting. Random assignment also supports a cause-and-effect interpretation for the volunteers studied, as discussed in Confidence Intervals for Differences in Experiments.

Common Mistakes and AP Exam Tips

  • Assuming that different labels mean independent groups: “Before” and “after” are different labels, but measurements from the same people are paired. Explain whether individuals contribute to one group or both.
  • Using equal sample sizes as evidence of independence: Two samples can have the same size and still be paired person by person. Sample size does not determine the design.
  • Ignoring matching: Different people can still form paired data if the study deliberately matches one person to another. The pair links their outcomes.
  • Assuming random assignment always creates independent samples: Random assignment to two separate groups can support an independent-groups comparison. But if each participant receives both treatments or is measured twice, the data are paired despite randomization of treatment order.
  • Applying the formula because the proportions are available: Being able to calculate two sample proportions does not establish that the interval formula is appropriate. Check the relationship between observations first.
  • Confusing independence with generalizability or causation: Random sampling and random assignment serve different purposes. Use the study design to explain what population or cause-and-effect claim the results may support.
AP Exam Tip: Name the observational units and state whether each one appears in one group or both. For independent groups, explain how the samples were selected or the units were assigned. For paired data, explicitly identify the repeated measurements or matched units and say that the independent two-proportion interval is not appropriate.

Key Takeaway

The two-proportion confidence interval is designed for a difference between proportions from independent groups. Separate random samples and separately assigned treatment groups can meet that design requirement when the other conditions are also checked. Repeated measurements on the same people and deliberately matched observations are paired, so their link must not be ignored or treated as if it were independence.

Key takeaway: Decide whether the observations are independent before calculating a two-proportion interval. If each unit contributes to both groups, or if observations are matched, use a method that accounts for the pairing rather than the independent-groups formula.

Check Your Understanding

For each scenario, decide whether an independent two-proportion interval is appropriate and explain the design reason.

  1. A random sample of 150 riders from one city and a separate random sample of 130 riders from another city are asked whether they wear a helmet. What design feature supports or challenges independence?
  2. A teacher records whether each of 40 students completes an assignment before and after a new reminder system. Are these two independent samples?
  3. In an experiment, participants are randomly assigned to use either a paper planner or a phone app, and each participant uses only one. What makes this different from a matched-pairs design?
  4. Researchers match each participant with a second participant of similar age, then assign one member of each pair to each of two programs. Are the two groups linked? Explain.
  5. Why is it not enough to calculate two sample proportions before deciding whether the two-proportion interval applies?