Tutorials › AP Statistics › Why Random Assignment Is the Key to Causation

Experimental design · Tutorial 183 of 1000

Why Random Assignment Is the Key to Causation

Understand how random assignment helps separate a treatment’s effect from the effects of other variables.

Beginner 9 min read

What You'll Learn

  • Explain how random assignment can distribute lurking variables across treatment groups.
  • Distinguish random assignment from random selection and state what each supports.
  • Describe why randomization does not guarantee identical groups.
  • Evaluate whether an observed group difference could be explained by a lurking variable.
  • Write a cautious cause-and-effect conclusion based on an experiment’s assignment method.

Why Assignment by Chance Matters

In Explanatory and Response Variables in Experiments, you learned to identify the condition researchers assign and the outcome they measure. But even when those roles are clear, another question matters: might some other variable help explain why the groups have different outcomes?

Suppose one group receives a new study app and another does not. If the app group also happens to include more students who already study regularly, a later score difference could reflect the app, prior study habits, or both. Prior study habits would be a possible lurking variable: a variable not included in the main comparison that may help explain an observed relationship. As explained in Confounding Variables Explained, a lurking variable can be confounding when it is related to both the explanatory variable and the response.

In an experiment, researchers can use chance to decide which experimental units receive each treatment. This is random assignment. Rather than letting participants choose a treatment or assigning them based on a researcher’s preference, researchers use a chance process, such as a random number generator.

Definition: Random assignment uses chance to allocate experimental units to treatment groups. It helps make the groups comparable, on average, with respect to potential lurking variables, so a difference in responses is less likely to be explained by a systematic pre-existing difference between the groups.

The key phrase is on average. Random assignment does not make every treatment group identical. It does not ensure that each group has precisely the same ages, prior experience, health, or other characteristics. Instead, it avoids a systematic assignment rule that would regularly put particular kinds of units into one group. Differences that remain can occur by chance.

How Randomization Helps Balance Lurking Variables

Imagine that some participants have more experience with a task than others. If researchers assign the most experienced participants to one treatment, experience may be mixed up with the treatment effect. If researchers assign participants by chance, each person has a chance of receiving each treatment. Experienced and inexperienced participants are not deliberately directed into particular groups.

The same logic applies to characteristics researchers did not measure or even think to record. Random assignment can help distribute both known and unknown potential lurking variables across the groups. That is an important advantage: researchers do not have to identify and control every possible influence individually for randomization to help.

Chance does not guarantee perfect balance in a particular experiment. A relatively small group could, just by chance, contain more experienced participants, older participants, or units with another relevant characteristic. Random assignment reduces the risk of a systematic imbalance; it cannot promise that no imbalance occurs.

Key distinction: Random assignment makes treatment groups comparable in expectation across repeated assignments, not necessarily identical in the one assignment that actually occurs. Any chance imbalance is still possible and should be considered when interpreting results.

When an experiment is conducted appropriately, a response difference between randomly assigned groups supports a cause-and-effect conclusion about the treatment for the experimental units in the study. Random assignment gives researchers a basis for attributing a group difference to the treatment rather than to a pre-existing difference that was systematically built into the groups. It does not eliminate every other concern, such as inconsistent treatment delivery, measurement problems, or participants dropping out unequally.

Random Assignment Is Not Random Selection

Random assignment and random selection both use chance, but they answer different questions. Random selection is about who enters a study from a population. Random assignment is about which treatment the study participants receive.

As discussed in Why Random Selection Matters and Scope of Inference: Four Combinations, random selection can support generalizing results to the population from which the sample was selected. Random assignment supports cause-and-effect conclusions about the treatments. A study can use one, both, or neither.

Chance processWhat is chosen by chance?What it can support
Random selectionIndividuals chosen for the studyGeneralizing from the sample to the population sampled from
Random assignmentTreatment received by each study participantCause-and-effect conclusions about the treatments

For example, volunteers might be randomly assigned to two treatments without being randomly selected from a larger population. That assignment helps compare the treatments fairly for those volunteers, but it does not by itself make the volunteers representative of everyone. Conversely, a random sample in an observational study can represent a population well while still not establishing that one observed condition caused another.

Worked Example: A Study App and Quiz Scores

Worked Example: Randomly Assigning Students to an App

A fictional school study includes 40 students who agree to test a study app. Researchers use a random number generator to assign 20 students to use the app for three weeks and 20 students to follow their usual study routine. Before the assignment, the students report their usual weekly study time. The app group’s mean is 6.1 hours, and the usual-routine group’s mean is 5.8 hours. At the end, the app group’s mean quiz score is 84.2 points and the usual-routine group’s mean is 78.6 points.

State: The question is whether assignment to the study app affects quiz scores for these 40 participating students. The explanatory variable is assigned study condition, with app and usual routine as its two levels. The response is quiz score in points.

Plan: Because the treatment conditions were assigned by chance, prior study habits were not used to place students into groups. Students who study more or less may still be unevenly distributed, but no deliberate assignment rule makes that happen. Random assignment helps account for prior study time and other potential lurking variables.

Do: The observed difference in mean quiz scores is \(84.2-78.6=5.6\) points. The app group’s mean prior study time was 0.3 hours higher: \(6.1-5.8=0.3\) hours. That small difference shows that random assignment did not produce perfectly identical groups on this recorded characteristic. It does not show that assignment failed; a chance process can leave group differences.

Conclude: The app group had a mean quiz score 5.6 points higher than the usual-routine group in this fictional study. Because students were randomly assigned, the design supports a cause-and-effect conclusion that using the app increased mean quiz scores for these participants, provided the study was carried out as described. The conclusion should not claim that random assignment made the groups identical or that the result necessarily applies to all students.

Worked Example: A Chance Imbalance in a Health Study

Worked Example: Comparing Two Symptom Treatments

In a fictional experiment, 48 adult volunteers are randomly assigned in equal numbers to a new symptom treatment or a comparison treatment. There are 24 volunteers younger than 40 and 24 who are at least 40. After assignment, the new-treatment group contains 13 volunteers younger than 40 and 11 who are at least 40. The comparison group contains 11 younger volunteers and 13 who are at least 40. By the end of the study, symptoms improved for 15 of the 24 volunteers receiving the new treatment and 10 of the 24 receiving the comparison treatment.

Identify the potential lurking variable: Age could be a lurking variable if it is related to symptom improvement. The researchers have recorded age group, so they can see how it was distributed, but recording it does not make the groups match exactly.

Compare the groups: The new-treatment group has two more younger volunteers than the comparison group: \(13-11=2\). The comparison group has two more volunteers aged 40 or older. This is an actual difference in the assigned groups, despite random assignment.

Describe the response difference: The improvement proportion is \(15/24=0.625\), or 62.5%, for the new treatment. For the comparison treatment it is \(10/24\approx0.4167\), or 41.7%. The observed difference is \(0.625-0.4167\approx0.2083\), which is about 20.8 percentage points.

Interpret carefully: The new-treatment group had a higher observed improvement proportion. Random assignment supports investigating this as an effect of the treatment rather than a systematic difference in how participants were assigned. However, age was not perfectly balanced, and random assignment alone does not prove that age or every other potential lurking variable was balanced in these 48 people. The conclusion should reflect the experiment’s results and its limits, not claim that chance removed all differences.

Worked Example: Seedling Growth and Greenhouse Position

Worked Example: Randomly Assigning a Fertilizer

A fictional greenhouse team assigns 30 similar seedlings at random: 15 receive a new fertilizer and 15 receive the standard growing mix. The team knows that one side of the greenhouse is usually warmer. After assignment, 9 of the new-fertilizer seedlings and 6 of the standard-mix seedlings are placed on the warmer side. After four weeks, the mean growth is 4.8 centimeters for the new-fertilizer group and 3.9 centimeters for the standard-mix group.

Find the possible lurking variable: Greenhouse position may affect growth because one side is warmer. If the warmer position is associated with greater growth, it could help explain differences in the response.

Assess the assignment: The team assigned fertilizer by chance, so it did not intentionally give the new fertilizer to seedlings on the warmer side. Still, the final placement counts are 9 versus 6 on that side. Chance has produced some imbalance in this particular assignment.

Compare the responses: The observed difference in mean growth is \(4.8-3.9=0.9\) centimeters in favor of the new fertilizer. The position imbalance means the team should be thoughtful when describing this result: some of the difference could be related to the warmer locations rather than the fertilizer.

Conclude: Random assignment helps make a treatment comparison fair by preventing researchers from systematically assigning seedlings to fertilizer based on position. It does not guarantee equal numbers in each position, especially in a study with only 30 seedlings. The result supports a causal interpretation of the fertilizer comparison, while the possible chance imbalance is a reason not to overstate certainty about the size of the fertilizer effect.

Common Mistakes and AP Exam Tips

  • Claiming random assignment guarantees balance. It does not guarantee equal numbers of every type of participant in every group. Say that it tends to balance potential lurking variables across groups over repeated assignments, while allowing for chance imbalances in a particular study.
  • Mixing up assignment and selection. Random assignment concerns treatment groups and supports cause-and-effect conclusions. Random selection concerns how participants are chosen and supports generalization to a population. Name the chance process the study actually used.
  • Saying randomization removes all lurking variables. Random assignment does not erase differences among individuals or prove that every relevant variable is balanced. It helps avoid systematic differences between treatment groups.
  • Treating any group difference as proof of a treatment effect. First describe the observed difference and the study design. Then explain why random assignment supports a causal conclusion, without claiming that chance guarantees certainty.
  • Generalizing to a broad population just because treatment assignment was random. Random assignment does not make study participants representative. Check whether the study also used random selection before making a population claim.
  • Ignoring the context in a conclusion. A complete response names the treatment, response, and study participants. For instance: “Among the participating students, random assignment supports concluding that using the app caused a difference in mean quiz scores.”

For full credit, be explicit about what chance did and what it did not do. State that random assignment helps create comparable treatment groups by distributing potential lurking variables, including ones researchers may not have measured. Then acknowledge that groups can still differ by chance and that random assignment alone does not justify generalizing beyond the study participants.

Key takeaway: Random assignment uses chance to allocate experimental units to treatments. It helps balance potential lurking variables across treatment groups, supporting cause-and-effect conclusions, but does not guarantee identical groups. Random selection and random assignment serve different purposes.

Check Your Understanding

Use the study’s assignment method and context to explain what its design can support.

  1. A fictional experiment randomly assigns 50 participants to two exercise plans. Explain how random assignment could help balance motivation, a potential lurking variable, and why exact balance is not guaranteed.
  2. A study randomly selects 200 residents for a survey but lets each resident choose whether to use a new health app. Does the study use random assignment? What kind of conclusion is limited by the absence of random assignment?
  3. Researchers randomly assign 36 devices to two battery settings. One group ends up with more devices from a particular production batch. Explain whether this proves random assignment was not used.
  4. A study uses volunteers and randomly assigns half to a new treatment and half to a comparison treatment. State what random assignment supports and what it does not establish about generalizing to all people.
  5. In a randomized plant experiment, the treatment group grows an average of 2 centimeters more than the comparison group, but the treatment group also has more plants in a sunny location. Explain how to describe the result without claiming perfect balance.