Tutorials › AP Statistics › Two-Sample Independence Conditions for Proportions

Two-proportion confidence intervals · Tutorial 505 of 1000

Two-Sample Independence Conditions for Proportions

Check whether the study design supports treating two groups and the observations within them as independent for a two-proportion interval.

Intermediate 10 min read

What You'll Learn

  • Distinguish random samples from random assignment and identify what each allows you to conclude.
  • Check whether the two groups were formed independently or consist of separate, nonoverlapping units.
  • Use the 10% condition to assess independence when sampling without replacement.
  • Recognize when repeated measurements, matched pairs, or clustering make observations dependent.
  • Explain the limits of a two-proportion interval when a design condition is not met.

Why the Study Design Comes First

In Standard Error for a Difference in Proportions, you used the two groups’ sample proportions to estimate the variability of \(\hat{p}_1-\hat{p}_2\). That calculation assumes a suitable study design. Before using a two-proportion confidence interval, check how the observations were selected or assigned and whether the groups and observations can reasonably be treated as independent.

There are two common ways to get a chance-based design: select random samples from populations or randomly assign experimental units to groups. These are not interchangeable. Random sampling can support generalizing results to the population sampled; random assignment can support a cause-and-effect conclusion about the units in the experiment. A study may use one, both, or neither.

Definition: For a two-proportion interval, the independence conditions concern both the relationship between the two groups and the relationship among observations within each group. Check whether the data come from appropriate random samples or random assignment, whether the groups are separate, and—when sampling without replacement—whether the 10% condition is met.

A Design Checklist for Two Groups

Start by identifying the units and the two groups. For example, the units might be households, patients, or seedlings. Then ask how the units entered the study and how they came to belong to Group 1 or Group 2. The group labels alone do not establish that the data are independent.

1
Check for a chance-based process.
Were the groups selected using random samples, or were experimental units randomly assigned to groups? State which process was used. Convenience samples or self-selected groups do not meet this condition simply because they are large.
2
Check independence between groups.
Are Group 1 and Group 2 made up of separate units, or were the groups sampled independently? If the same units contribute to both groups, the observations are paired rather than independent.
3
Check independence within each group.
Consider how observations were selected and whether one unit’s result could affect another’s. When sampling without replacement from a finite population, check the 10% condition separately for each group.

The 10% condition says that, when a random sample is drawn without replacement from a finite population, the sample size should be no more than 10% of that population. For Group \(i\), check \(n_i\leq 0.10N_i\), where \(n_i\) is its sample size and \(N_i\) is the population size. This makes it reasonable to treat observations within that sample as approximately independent.

$$ n_i\leq 0.10N_i \qquad\text{or equivalently}\qquad N_i\geq 10n_i $$

Apply the condition to each sampled population, not to the two sample sizes added together. If the two samples come from distinct populations, check each population against its own sample size. If groups are formed by random assignment in an experiment, the 10% condition for sampling without replacement is not the relevant check: the units were assigned, not sampled from a finite population for the purpose of estimating its proportion.

Independence also involves the structure of the data. Two samples from separate groups are not independent if each person in one sample is deliberately matched with a person in the other and the analysis treats the paired results as unrelated. Nor are repeated measurements on the same person independent observations. Likewise, observations from people in the same household or tightly connected cluster may influence one another. Look at the actual unit of observation and the way data were collected.

Random Samples and Random Assignment: Different Strengths

A random sample uses chance to select units from a population. If two random samples are selected separately from two populations, this supports inference about each population, provided the other relevant conditions are met. The sample groups should be separate, and each sample should satisfy its own 10% condition when drawn without replacement.

Random assignment uses chance to place experimental units into treatment groups. It supports a comparison of the treatments for the experimental units because the assignment process helps balance other factors across groups. To use a two-proportion interval in this setting, verify that each unit is assigned to only one group and that outcomes from different units can reasonably be treated as independent. Avoid interference: one unit’s treatment or outcome should not change another unit’s outcome.

Random assignment does not, by itself, make the experimental units a random sample from a wider population. For example, if volunteers are randomly assigned to two treatments, a difference can support a causal comparison for those volunteers, but it does not automatically represent all people who might use the treatment. As covered in Scope of Inference Based on How Data Were Collected, the way units are selected and the way treatments are assigned answer different questions.

Key distinction: Random sampling is about who is selected and can support generalizing to a population. Random assignment is about which treatment a unit receives and can support a cause-and-effect conclusion. Neither label excuses checking whether the observations are paired, clustered, or otherwise dependent.

Worked Examples

Worked Example: Two Random Samples of Households

Setting: Imagine a survey comparing whether households in two separate towns have a home compost bin. A computer selects a random sample of 240 households from Town 1, which has 3,600 households, and a separate random sample of 180 households from Town 2, which has 2,500 households. Each household is asked once. Check the chance-based design, independence between groups, and the 10% condition within each group.

Random samples: The households in each town were selected by a random process. The design therefore meets the random-sampling condition for the two towns.

Independence between groups: Town 1 and Town 2 are separate populations, and the samples were selected separately. A household cannot belong to both town samples, so the groups consist of distinct units.

Independence within groups: Each household contributes one response. Check the sample fractions:

$$ \frac{240}{3600}=0.0667 \qquad \frac{180}{2500}=0.072 $$

The first sample is about \(6.67\%\) of Town 1’s households, and the second is \(7.2\%\) of Town 2’s households. Each is less than \(10\%\); equivalently, \(3600\geq10(240)=2400\) and \(2500\geq10(180)=1800\). Thus, both samples satisfy the 10% condition.

Conclusion: The random-sampling condition, separate-group condition, and within-group 10% checks are met. These design checks support treating the groups and observations as independent for the two-proportion interval. They do not, on their own, check every condition needed for the interval.

Worked Example: Random Assignment in a Seedling Experiment

Setting: Imagine a greenhouse experiment with 120 seedlings. Researchers randomly assign 60 seedlings to a new nutrient solution and 60 to the usual solution. Each seedling grows in its own pot, receives only its assigned solution, and is classified once as having reached a specified growth milestone by the end of the study. Check whether the design supports a two-group comparison.

Random assignment: The researchers used chance to assign the seedlings to the two solutions. This meets the random-assignment condition for comparing the treatments.

Independence between groups: Each seedling is assigned to exactly one solution, so the treatment groups contain separate seedlings. Separate pots and careful treatment delivery also help prevent one seedling’s treatment from affecting another seedling’s outcome. Given this setup, it is reasonable to treat the groups as independent.

Independence within groups: Each seedling contributes one outcome, and each is grown in a separate pot. The setup gives no stated reason for one seedling’s outcome to determine another’s. The observations within each group can therefore reasonably be treated as independent for this comparison.

Conclusion: Random assignment, separate treatment groups, and a setup that limits interaction between seedlings support the independence conditions. The 10% condition is not the relevant check here because the seedlings were assigned to treatments rather than selected as samples without replacement from a population. Since the seedlings were not described as a random sample of all seedlings, random assignment alone does not justify generalizing the results to every seedling.

Worked Example: Random Samples That Fail the 10% Condition

Setting: Imagine a public library system with 800 registered members. A researcher takes a random sample of 120 members and records whether each member uses an e-reader. A second, separate library system has 700 members; the researcher takes a random sample of 100 members and records the same characteristic. Check the random-sampling and independence conditions.

Random samples and separate groups: Both samples are selected at random, and they come from separate library systems. The groups therefore meet the random-sampling and between-group checks.

Check the 10% condition for each sample:

$$ \frac{120}{800}=0.15 \qquad \frac{100}{700}\approx0.1429 $$

The first sample is \(15\%\) of its population, and the second is about \(14.29\%\). Both exceed \(10\%\). In the equivalent form, \(800<10(120)=1200\), and \(700<10(100)=1000\). So neither sample meets the 10% condition.

Conclusion: Random selection and separate populations do not repair the failed within-group checks. Under the usual AP condition check, the two samples cannot be treated as approximately independent within groups on the basis of the 10% condition. The researcher should report this limitation rather than claim all conditions are met; the standard two-proportion interval is not justified by the stated design conditions.

Worked Example: The Same People Measured Twice

Setting: Imagine 75 library members answer whether they use an audiobook app before and after a one-month trial. The researcher wants to compare the proportion answering “yes” at the two times and proposes an ordinary two-proportion interval, calling the before group Group 1 and the after group Group 2. Are the two groups independent?

Random process: Suppose the 75 members were selected at random from a large membership list. That supports the sample-selection condition, and \(75\) is no more than \(10\%\) of the membership if the list contains at least \(750\) members.

Independence between groups: The before and after observations are not from separate groups of people. Each member appears in both sets of observations. A person’s answers across time may be related, so the two groups are paired, not independent.

Conclusion: Even with random selection and a satisfied 10% condition, the ordinary two-proportion interval’s between-group independence condition is not met. The issue is not the size of the samples; it is that the same people supply both measurements. The proposed independent two-group method is inappropriate for these data.

Common Mistakes and AP Exam Tip

  • Writing only “the data are random”: Say whether there were random samples or random assignment, and name the groups or units involved.
  • Checking only between-group independence: Separate groups do not guarantee independent observations within each group. Check repeated responses, clustering, and the 10% condition when applicable.
  • Combining sample sizes for the 10% condition: Check \(n_i\leq0.10N_i\) separately for each population sampled without replacement.
  • Applying the 10% condition to the wrong process: The condition is for sampling without replacement from a finite population. Random assignment is a different design step.
  • Treating repeated measurements as independent groups: If the same units contribute to both proportions, identify the paired structure instead of treating the samples as separate.
  • Claiming random assignment allows generalization: Assignment supports cause-and-effect reasoning; random selection supports generalizing to a population. Check which process the study actually used.
AP Exam Tip: A full-credit condition statement identifies the sampling or assignment process, explains why the groups are separate or paired, and checks the 10% condition for each sample drawn without replacement. Use the actual population and sample sizes in your check, then state clearly whether each condition is met.

Key Takeaway

Before using a two-proportion interval, inspect the study design rather than relying on the group labels or the standard-error calculation. Establish whether chance was used to select samples or assign treatments, check that the groups are genuinely separate, and consider whether observations within each group are independent. When a random sample is drawn without replacement, verify the 10% condition for that group.

Key takeaway: Random samples or random assignment provide the chance-based design; separate groups support independence between groups; and the 10% condition supports approximate independence within each sample drawn without replacement. Repeated measurements and paired data are not independent two-group samples.

Check Your Understanding

For each situation, identify the relevant design checks and state whether the two-group independence conditions are supported.

  1. A random sample of 90 residents is taken from a town of 1,500, and a separate random sample of 110 is taken from a neighboring town of 2,000. Check the 10% condition for both samples.
  2. In an experiment, 48 volunteers are randomly assigned to one of two app designs. Each volunteer uses only one design. What does random assignment support, and what does it not establish about generalizing to all app users?
  3. A researcher compares a yes/no response from the same 40 students before and after a workshop. Which independence check fails for an ordinary two-proportion interval?
  4. Why should the 10% condition be checked separately for each group when two random samples come from different populations?
  5. Give one example of how observations within a group might fail to be independent even when the group was selected at random.