Tutorials › AP Statistics › Experiments and Establishing Causation

Regression and context · Tutorial 927 of 1000

Experiments and Establishing Causation

Learn how to design and evaluate a randomized experiment that can support a cause-and-effect conclusion without confusing random assignment with random sampling.

Intermediate 9 min read

What You'll Learn

  • Distinguish an experiment from an observational study by how the explanatory condition is assigned
  • Explain how random assignment supports a cause-and-effect conclusion
  • Identify the experimental units, treatments, and response in a study
  • Use comparison, control, and replication to plan a fair experiment
  • Distinguish random assignment from random sampling and describe what each allows
  • Recognize why randomization supports, but does not guarantee, a causal conclusion

From an Association to a Cause-and-Effect Question

In “Lurking Variables in Regression Settings,” you considered how an unmodeled variable can contribute to an observed association. This raises a practical question: how could a study provide stronger evidence that changing an explanatory variable actually changes a response? The key is to control how the explanatory condition is assigned.

In an observational study, researchers record what happens without assigning the explanatory condition to the individuals or cases. In an experiment, researchers deliberately impose one or more conditions, called treatments, and measure the response. When treatments are assigned to experimental units at random, a well-designed experiment can provide convincing evidence about whether a treatment caused a difference in the response.

Definition: An experiment is a study in which researchers deliberately impose treatments on experimental units and measure responses. Random assignment uses chance to allocate the experimental units to treatment groups. A well-designed randomized experiment provides a basis for a cause-and-effect conclusion about the treatments in the study.

An experimental unit is the person, animal, object, or other unit that receives a treatment. When the units are people, they are often called subjects or participants. The treatment is the condition imposed, and the response variable is the outcome measured after the treatment. In an experiment comparing two study routines, for example, students are the experimental units, the routines are treatments, and a later quiz score could be the response.

The distinction matters because an association alone does not show that changing one variable causes a change in another. As discussed in “Association Versus Causation in Regression” and “Observational Studies and Confounding,” people who choose different conditions may also differ in other ways. Random assignment helps make the treatment groups comparable, on average, with respect to both measured and unmeasured characteristics.

Why Random Assignment Supports Causation

Suppose participants choose whether to use a new productivity app. Those who choose it may already be more organized or motivated than those who do not. If their work output later differs, the app may not be the only explanation. This is the kind of concern raised by confounding: preexisting differences may be mixed together with the effect of the explanatory condition.

In a randomized experiment, chance—not a participant’s preference or a researcher’s judgment—determines who receives each treatment. Random assignment does not make every group exactly alike. By chance, one group might still include more experienced participants. But a fair random process tends to distribute relevant characteristics across groups rather than systematically assigning them to one treatment. This makes treatment the main planned difference between the groups.

Key idea: Random assignment helps create comparable treatment groups before the response is measured. If the groups are treated alike in other important ways and differ in their assigned treatment, a response difference provides evidence that the treatment caused a difference in the study.

Randomization does not eliminate all uncertainty. Chance differences between groups can remain, and a response difference might occur by chance even if the treatments have no real effect. The conclusion should therefore be based on the study’s results and acknowledge uncertainty. In later inference topics, statistical procedures will help assess how compatible the observed difference is with chance variation.

Random assignment is also different from random sampling. Random assignment determines which treatment a study participant receives; random sampling determines which individuals from a population enter the study. Random assignment supports causal conclusions for the experimental units studied. Random sampling supports generalizing results to a larger population. One process does not automatically provide the other.

Design featureWhat chance determinesWhat it supports
Random assignmentWhich treatment each experimental unit receivesCause-and-effect conclusions about the treatments
Random samplingWhich individuals are selected from a populationGeneralizing results to that population

Principles of a Well-Designed Experiment

A strong experiment uses comparison, random assignment, and replication. These principles work together. Comparison gives the results a reference point; random assignment reduces systematic differences among treatment groups; and replication means applying each treatment to enough experimental units to see whether a pattern is consistent rather than dependent on a few unusual cases.

  • Comparison: Compare the response under at least two treatments or conditions. A control group provides a useful reference, such as the usual procedure, a placebo, or no active treatment when appropriate.
  • Random assignment: Use a chance process to assign experimental units to treatments. Do not let participants choose their treatment or assign treatments based on a characteristic that could affect the response.
  • Control: Keep other aspects of the experiment as similar as practical across treatment groups. This includes procedures, timing, instructions, and measurement methods, so the assigned treatment is the main planned difference.
  • Replication: Assign enough experimental units to each treatment to make results less dependent on the outcomes of just a few units. Repeating measurements on one person is not the same as assigning multiple independent people to a treatment.

A placebo is an inactive treatment designed to resemble an active treatment. When participants do not know which treatment they receive, they are blinded. If the people measuring or evaluating responses also do not know the assignments, the study can be double-blind. Blinding can reduce the influence of expectations or biased measurements, but it does not replace random assignment. Some treatments cannot be concealed, and an experiment can still use random assignment and careful, consistent measurement.

1
State the question and response.
Specify what effect is being studied and how the response will be measured.
2
Identify units and treatments.
Name the experimental units and the conditions that will be imposed on them.
3
Assign treatments by chance.
Use a random method so each unit’s treatment is not determined by choice or a researcher’s preference.
4
Keep procedures comparable and replicate.
Apply the treatments consistently, measure the response in the same way, and include enough units in each group.
5
Compare responses and limit the conclusion appropriately.
Describe the treatment-group difference in context. Use causal wording for the assigned treatments, and generalize only when the sampling design supports it.

Worked Examples: Evaluating Causal Evidence

Worked Example: A Reminder Message and Appointment Attendance

A fictional clinic wants to know whether a text reminder sent two days before an appointment increases attendance. The clinic identifies 120 patients with upcoming appointments who have agreed to participate. It randomly assigns 60 patients to receive the reminder and 60 to receive the clinic’s usual appointment information. The clinic records whether each patient attends.

State. The question is whether receiving the text reminder causes a change in appointment attendance among the participating patients. The response is whether a patient attends the scheduled appointment.

Plan and check the design. The experimental units are the 120 patients. The two treatments are receiving the extra text reminder and receiving the usual information. The clinic uses random assignment to place 60 patients in each group, so participants do not choose their treatment and staff do not assign it based on a prediction of who is likely to attend. The same attendance definition and tracking procedure should be used for both groups. With 60 patients assigned to each condition, the experiment has replication in both groups.

Do. The clinic compares the proportion who attend in the reminder group with the proportion who attend in the usual-information group. The comparison is meaningful because the main planned difference between the groups is the added reminder. Random assignment helps distribute other influences on attendance, such as work schedules or transportation access, across the groups.

Conclude. If attendance is higher in the reminder group, the randomized design supports the conclusion that the reminder caused an increase in attendance for the participating patients, subject to chance variation and the study’s execution. The clinic should not automatically claim the result applies to all patients everywhere: the participants were not described as a random sample from a broader population.

Worked Example: Comparing Two Seed Treatments

A fictional greenhouse tests whether a seed coating affects the height of young tomato plants. Staff place 80 similar seedlings in separate pots. They label 40 pots for a coating treatment and 40 for no coating, then use a random-number generator to assign the seedlings to the two groups. All plants receive the same soil, light, water schedule, and growing period. At the end, staff measure plant height in centimeters.

Identify the design. The experimental units are the 80 seedlings. The treatments are the seed coating and no coating. Height in centimeters is the response. This is an experiment because staff deliberately impose the coating condition rather than simply observing which plants happened to receive it.

Explain the role of random assignment. Random allocation helps prevent staff from placing the strongest-looking seedlings in one group. It also tends to distribute other differences among seedlings across both treatments. The seedlings may not be identical after random assignment, but the procedure avoids a systematic choice that could favor one treatment.

Explain comparison, control, and replication. The uncoated group is a comparison condition. Using the same growing conditions controls other planned influences on height. Applying each treatment to 40 seedlings provides replication, so the comparison is not determined by only one or two plants.

Conclude cautiously. If the coated seedlings tend to be taller, the design supports a cause-and-effect conclusion about the coating under these greenhouse conditions. It does not by itself establish that the same effect will occur for every tomato variety or in outdoor growing conditions.

Worked Example: Why Random Sampling Is Not Random Assignment

A fictional school randomly selects 100 students to answer a survey about nightly sleep and quiz performance. The students report their usual sleep, and the school compares quiz scores for students who report more sleep with scores for those who report less.

Identify the study type. This is an observational study. The school records students’ existing sleep habits; it does not assign students to sleep for a specified number of hours. Randomly selecting students for the survey is random sampling, not random assignment.

Describe what the design can support. The survey can describe an association between reported sleep and quiz scores among the sampled students. It cannot establish that increasing sleep causes quiz scores to rise. Students with different sleep habits may also differ in stress, study time, health, or other factors related to quiz performance.

Consider an experiment. If it were ethical and practical, researchers might randomly assign volunteers to different, reasonable sleep-schedule recommendations and measure a consistent outcome afterward. Random assignment could support a causal conclusion about those assigned recommendations. It would not necessarily make the experiment a random sample of all students, so generalization would still depend on how participants were recruited.

Conclude. The survey’s random sample can support generalization to the population from which students were sampled, subject to the survey’s limitations. It does not turn the observational association into evidence of causation.

Limits, Ethics, and AP Exam Wording

Randomized experiments are not always possible or ethical. Researchers cannot randomly assign people to harmful exposures simply to test whether those exposures cause damage. In such settings, observational evidence can still be informative, but possible confounding limits how confidently a cause-and-effect conclusion can be made. As in the earlier tutorials on confounding and lurking variables, name plausible alternative explanations rather than treating an association as proof of cause.

Even a carefully randomized experiment can be weakened by poor execution. Participants might not follow their assigned treatment, some responses might go unrecorded, or the measurement method might favor one group. Researchers should describe these issues and avoid claiming that randomization guarantees a perfect comparison. Random assignment makes causal interpretation more defensible; it does not remove every source of bias or chance variation.

  • Do not call a study an experiment just because it compares groups. Researchers must impose treatments for it to be an experiment.
  • Do not confuse random selection and random assignment. Sampling concerns who enters a study; assignment concerns which treatment those participants receive.
  • Do not say random assignment guarantees identical groups. State that it tends to create comparable groups by distributing characteristics through chance.
  • Do not claim a causal result proves a universal effect. A randomized experiment supports causation for the treatment comparison studied; generalization requires appropriate sampling and careful attention to the study setting.
  • Do not omit the response or context. A full-credit explanation names what was assigned, what was measured, and which units received the treatments.

On an AP-style question, distinguish the design first, then connect the design to what can be concluded. A strong response might say: “Because the researchers randomly assigned the participants to the two treatments and measured the same response in each group, a difference in the responses provides evidence that the assigned treatment caused a difference for these participants. Since the participants were volunteers rather than a random sample, the result cannot automatically be generalized to all members of the population.”

Key takeaway: A well-designed randomized experiment uses chance to assign treatments, compares groups under consistent conditions, and includes replication. Random assignment supports cause-and-effect conclusions about the treatments studied; random sampling, not random assignment, is what supports generalization to a population.

Check Your Understanding

For each situation, identify the design and state what conclusion it can support.

  1. A researcher assigns volunteers by coin toss to use one of two practice schedules, then compares their performance on the same skills test. What feature supports a causal conclusion?
  2. A town randomly selects residents for a survey about cycling and fitness. Residents report their own cycling habits. Is this random assignment, random sampling, or both? What can the study conclude?
  3. A greenhouse assigns one treatment to its tallest seedlings and another treatment to its shortest seedlings. Explain why a later difference in growth would be difficult to attribute to treatment.
  4. In a randomized experiment, treatment groups differ in their average response. Why is it still important to acknowledge chance variation?
  5. Write one sentence explaining how a randomized experiment can support causation while a random sample can support generalization.