Tutorials › AP Statistics › Checking Independence Between Two Groups

Conditions for mean inference · Tutorial 653 of 1000

Checking Independence Between Two Groups

Use the people, objects, and selection or assignment process—not the group labels or sample sizes—to decide whether two groups are independent or paired.

Intermediate 8 min read

What You'll Learn

  • Distinguish separate independent samples from measurements or observations that form pairs.
  • Use a linkage audit to check whether the same units or deliberately matched units appear in both groups.
  • Explain why groups drawn from separate populations may reasonably be treated as independent.
  • Recognize why disjoint groups are not automatically independent when their selection process links them.
  • Write a context-specific independence justification for a two-sample t procedure.

Two Groups Are Not Necessarily Independent

In “Conditions for Two-Sample t Procedures,” we checked independence within each group and independence between groups. This tutorial focuses on the second check: whether the design makes the two groups reasonably independent, or instead connects their observations in pairs.

The distinction is about how the data were produced. Two columns in a table do not automatically represent independent samples. The same participant might contribute an observation to each column, or different participants might have been deliberately matched. In either case, the observations are linked, and a paired t procedure—not a two-sample t procedure—is generally appropriate. By contrast, two samples made up of different, unmatched units can often be treated as independent if the sampling or assignment design supports that conclusion.

Definition: Two groups are independent for a two-sample t procedure when observations in one group are not linked to observations in the other group through shared units, deliberate matching, or a selection or assignment design that makes the groups dependent. Paired data have a meaningful one-to-one link between observations, such as repeated measurements on the same unit or measurements on matched units.

“Independent” does not mean that the two groups have the same sample size, similar averages, or different labels. It is a claim about the design. To assess it, identify the observational units and trace whether an observation in one group has a specific partner in the other.

Use a Linkage Audit

A useful new technique is a linkage audit: inspect the design for ways that observations across the groups are connected. Begin with the unit that contributes one measurement. Then ask whether the same unit appears in both groups, whether different units were intentionally matched, and whether the way the groups were selected or assigned ties them together.

1
Name the observational unit.
State what one measurement represents: a person, tree, device, school, or another unit. Do not confuse the unit with the measurement itself.
2
Look for repeated units.
If the same unit contributes one measurement to each group—for example, before and after measurements—the observations are paired.
3
Look for deliberate matching.
Different units can still be paired if the study matched them in meaningful pairs, such as siblings or participants with similar starting measurements.
4
Inspect how groups were formed.
If there are no pairs, explain how distinct units were sampled or assigned and whether the design supports treating the groups as independent.

When observations are paired, each pair—not each individual measurement—is the basic comparison unit. In a paired t procedure, the analysis uses one difference per pair. As covered in “Conditions for One-Sample Versus Paired Data,” those differences must themselves meet the relevant conditions for paired inference. Do not switch to a two-sample t procedure just because the paired data are displayed in two columns.

When no meaningful pairing exists, a two-sample t procedure may be appropriate, subject to the other conditions from “Conditions for Two-Sample t Procedures.” Independence between groups is only one part of that assessment. Randomness, independence within each group, and the Normal/Large Sample condition for each group still need their own checks.

What Counts as Evidence of Independence?

A clear justification names the design features that separate the groups. For example, the study might take separate random samples from two distinct populations, with no unit eligible to appear in both samples. Or a randomized experiment might assign different participants to one of two treatments, with no matching or repeated measurements. Those features support treating the groups as independent.

If random samples are drawn without replacement from finite populations, the 10% condition helps support independence within a sample, as explained in “Checking the 10% Condition for Independence.” When the two samples come from distinct populations, check the condition for each sample against its own population. If both groups are selected without replacement from one shared finite population, the selection process itself can link the groups; do not claim independence just because the selected people are distinct. Describe the design and check the sampling fractions with care.

A study can also have distinct people in the two groups but still deliberately match them. For example, researchers may pair each participant receiving one treatment with a participant who has a similar starting measurement. The people differ, but each observation has a designated partner. That is paired data.

Conversely, two groups can be independent even if they use the same measurement method or are studied at the same time. What matters is whether the units or the process link observations across groups—not whether the groups share a setting, instrument, or schedule.

Design featureImplication for the comparison
Same units measured under both conditionsPaired data; compare within-unit differences.
Different units deliberately matched into pairsPaired data; compare the matched-pair differences.
Different, unmatched units selected from separate populationsIndependence between groups may be reasonable; explain the sampling design.
Different, unmatched participants randomly assigned to separate treatmentsIndependence between groups may be reasonable; explain the assignment design.
Distinct units, but a shared selection process may link which units enter each groupDo not assume independence from distinct identities alone; assess the selection design.

Worked Examples

Worked Example: The Same People Try Two Study Schedules

A fictional learning center asks 16 volunteers to complete a practice quiz after using a visual-review schedule and another practice quiz after using an audio-review schedule. Each volunteer tries both schedules. The center wants to compare the mean quiz scores under the two schedules.

State. Let the two groups be the 16 scores after visual review and the 16 scores after audio review. We need to decide whether those groups are independent or paired before selecting a mean procedure.

Plan. Use a linkage audit: identify the observational unit, check whether it appears in both groups, and decide whether each observation has a partner.

Do. The observational unit is a volunteer’s quiz score under a schedule. Each volunteer contributes one score to each group, so each visual-review score has a specific audio-review score from the same person as its partner. The groups therefore are paired, not independent. There are 16 pairs, not 32 unrelated participants.

Conclude. A two-sample t procedure would ignore the within-person links. A paired t procedure is the appropriate mean procedure to consider, using one difference between the two scores for each volunteer. Its conditions, including the shape of the differences, would still need to be checked.

Worked Example: Participants Are Matched but Not Reused

A fictional clinic compares two exercise plans. Before the study, it matches 20 participants in Plan A with 20 different participants in Plan B who have similar starting activity levels. Each participant follows only one plan, and the clinic measures weekly activity after six weeks.

State. Let the groups be the weekly activity measurements for participants assigned to Plan A and Plan B. We are assessing whether the groups are independent or paired.

Plan. Check for repeated units and for deliberate matching. Having different people in the groups would rule out repeated measurement on the same person, but it would not by itself rule out pairing.

Do. No participant follows both plans, so no person is measured in both groups. However, each Plan A participant was deliberately matched with a particular Plan B participant based on starting activity. That one-to-one design connects the measurements across groups.

Conclude. These are matched pairs, not independent samples. To compare mean activity under the plans, analyze one activity difference for each matched pair using a paired t procedure, after checking the conditions for those differences. The fact that the matched participants are different people does not erase the pairing.

Worked Example: Separate Random Samples From Separate Populations

A fictional parks department wants to compare mean trail-repair time, in hours, for two regions. It randomly selects 24 repair jobs from a large list of jobs in the north region and 22 different jobs from a large list in the south region. No job appears on both lists, and jobs are not matched. Each region has more than ten times as many eligible jobs as the number selected from it.

State. Let \(\mu_N\) be the mean repair time, in hours, for eligible north-region jobs, and let \(\mu_S\) be the corresponding mean for eligible south-region jobs. We are checking whether the groups can be treated as independent for a two-sample t procedure comparing \(\mu_N-\mu_S\).

Plan. Check the sampling design for a connection between groups, then check within-sample independence using the 10% condition. These checks address independence; the random and shape conditions from “Conditions for Two-Sample t Procedures” must also be considered before using the procedure.

Do. The observational unit is a repair job. The samples were selected from separate regional lists, and no job is in both groups or matched with a job in the other region. Thus, the design supports treating the groups as independent. Within the north sample, the population contains more than \(10(24)=240\) eligible jobs, so 24 is no more than 10% of that population. Within the south sample, there are more than \(10(22)=220\) eligible jobs, so 22 is no more than 10% of that population. The 10% condition supports independence among sampled jobs within each regional sample as well.

Conclude. The separate random samples, distinct unmatched jobs, and 10% checks support independence within and between the groups. The random selection supports generalizing to eligible jobs in the respective regions. This independence check does not by itself establish that every condition for a two-sample t procedure is satisfied; the shape of each group and the other conditions still need review.

Common Mistakes and AP Exam Tips

  • Calling groups independent because the people are different. Different people can be matched deliberately. Say whether matching occurred, not only whether the identities differ.
  • Calling every two-column data set independent. Before-and-after measurements are linked even when shown in separate columns. Identify whether the same unit contributes to both.
  • Assuming equal sample sizes mean pairing. Equal counts do not create pairs. Pairing requires a meaningful one-to-one link in the design.
  • Assuming unequal sample sizes rule out pairing. A study might have missing measurements in some pairs. Explain the design and identify the actual complete pairs rather than inferring the design from the counts.
  • Claiming “the groups are independent” without evidence. A stronger response names the distinct units, lack of matching or repeated measurements, and separate sampling or assignment process.
  • Stopping after the independence check. Independence is not the only condition for a two-sample t procedure. Continue to assess randomness, independence within each group, and the shape condition separately for each group.

A full-credit explanation is specific: “Different, unmatched repair jobs were randomly selected from separate regional lists, and no job appears in both samples, so the groups can reasonably be treated as independent. Each sample is also no more than 10% of its regional population, supporting independence within each sample.” If the same units were measured twice or units were deliberately matched, state that the data are paired and consider a paired t procedure instead.

Key takeaway: Independence between groups is determined by the design. Use a linkage audit to look for repeated units, deliberate matching, and selection or assignment links. Distinct group labels—or even distinct units—are not enough on their own; explain why the groups can reasonably be treated as independent, or recognize that the data are paired.

Check Your Understanding

For each situation, decide whether the groups are independent or paired, and identify the design evidence behind your decision.

  1. A group of 12 cyclists records its time on a route before and after a training plan. What type of data does the design produce?
  2. Researchers compare two treatments using different participants, but match each participant in one treatment group with a participant of similar age in the other. Are the groups independent?
  3. Two random samples are selected from separate populations, and no individual can appear in both. What design details support treating the groups as independent?
  4. Two groups contain 25 observations each. Does the equality of the sample sizes show that the observations are paired? Explain.
  5. Two groups contain different people, but were selected sequentially without replacement from one shared finite population. Why should the analyst examine the selection process before claiming independence between groups?