The Study Design Sets Limits on the Conclusion
In Observational Studies and Their Limits and Experiments and Why They Can Show Cause, you learned that studies differ in whether researchers assign treatments. The distinction matters when interpreting an apparent relationship. Two variables can be associated—values of one tend to occur alongside particular values of the other—without the data showing that one variable caused the other.
The central question is not simply, “Did the groups have different outcomes?” It is also, “How were the groups formed?” If people chose their own activities, or if researchers recorded conditions that already existed, the groups may differ in other ways. If researchers assigned treatments at random, that design gives stronger support for attributing a difference in outcomes to the treatments.
As in Confounding Variables Explained, another variable may help explain why groups differ. The key step in this tutorial is to use the study’s data-collection design to judge what kind of conclusion is justified. A careful conclusion describes what the design supports and avoids claiming more.
A Design-Based Decision Process
Start by identifying the explanatory variable and response. Then ask whether the researchers deliberately imposed different treatments. If they did, ask whether experimental units were randomly assigned to the treatments. If they did not assign treatments, the study is observational, even if the researchers followed participants over time or collected detailed records.
Name the explanatory variable or treatment and the response being measured.
Did researchers assign treatments, or did they observe existing choices, traits, or conditions?
If treatments were imposed, determine whether chance was used to assign experimental units to treatment groups.
Observational data support describing an association. A well-designed randomized experiment can support a cause-and-effect conclusion about the treatments compared.
In an observational study, researchers do not control which explanatory-variable values participants receive. People may select their own behavior, or circumstances may determine their group. Those groups can differ in other characteristics related to the response. A lurking or confounding variable is one possible explanation, but even if no particular confounder is obvious, the observational design does not rule out other explanations.
In an experiment, researchers impose treatments. When they also use random assignment, chance determines which experimental units receive each treatment. This tends to create groups that are similar, on average, in other characteristics—both measured and unmeasured. Differences in the response can then be attributed more plausibly to the assigned treatments rather than to pre-existing differences between groups. Random assignment does not guarantee that the groups will be perfectly alike, especially in a small experiment, but it helps control confounding.
“Can support” is important wording. Random assignment strengthens a causal argument; it is not a guarantee that every part of a study was carried out perfectly or that chance could not contribute to the observed difference. The conclusion should remain connected to the treatments, response, and experimental units actually studied.
Random Assignment Is Not Random Selection
Random assignment and random selection answer different questions. Random assignment uses chance to place experimental units into treatment groups. It helps make the groups comparable and supports cause-and-effect conclusions. Random selection uses chance to choose participants from a population. It helps make a sample representative and supports generalizing results to that population, when the sampling process is appropriate.
A study can use random assignment without randomly selecting its participants. In that case, the experiment may support a cause-and-effect conclusion for the units studied, but the results may not automatically generalize to a broader population. Conversely, a random sample in an observational study may help represent a population, but it does not make the observed relationship causal. The next tutorial, Generalizing Results to a Population, develops the separate question of who the results can represent.
Worked Example: Study-App Use and Exam Results
Worked Example: Study-App Use and Exam Results
A school counselor reviews app records and exam results for students who chose whether to use a study app. Among 120 students who used the app, 48 passed the exam. Among 100 students who did not use it, 30 passed. A report says, “Using the app caused students to pass.” Is that conclusion justified?
The explanatory variable is whether a student used the app, and the response is whether the student passed. The counselor did not assign students to use or not use the app; students chose for themselves. This is an observational study.
The pass rate among app users is \(48/120=0.40\), or 40%. Among nonusers, it is \(30/100=0.30\), or 30%. In these records, app users had a pass rate 10 percentage points higher than nonusers. That is a descriptive association in the students studied.
The causal claim is not justified by this design. Students who chose to use the app might differ in prior preparation, time spent studying, or other characteristics connected to passing. These are possible alternative explanations, not proof of what caused the difference. The data do not show that app use caused the higher pass rate.
A design-based conclusion would be: “In the students whose records were reviewed, app use was associated with a higher pass rate: 40% for users and 30% for nonusers. Because students chose whether to use the app, these data do not establish that using the app caused the difference.”
Worked Example: A Randomized Reminder Experiment
Worked Example: A Randomized Reminder Experiment
A fictional research team wants to know whether a text reminder increases daily water intake among volunteers in a summer program. The team randomly assigns 100 volunteers to receive a reminder and 100 to receive no reminder. At the end of the study, 80 volunteers in the reminder group and 62 in the no-reminder group meet the program’s hydration goal. What conclusion is justified?
The explanatory variable is the assigned condition: receiving a text reminder or receiving no reminder. The response is whether a volunteer meets the hydration goal. Researchers deliberately imposed the two conditions, so this is an experiment. The volunteers were randomly assigned to the conditions.
The percentage meeting the goal in the reminder group is \(80/100=0.80\), or 80%. In the no-reminder group, it is \(62/100=0.62\), or 62%. The observed difference is \(80\%-62\%=18\) percentage points, with the higher percentage in the reminder group.
Because the experiment used random assignment, it supports the conclusion that assignment to the reminder condition caused a higher rate of meeting the hydration goal among these volunteers under the study conditions. Random assignment makes pre-existing differences a less convincing explanation for the observed group difference, though chance variation or problems in carrying out the study could still matter.
A careful conclusion is: “Among the volunteers in this experiment, assignment to receive a text reminder caused a higher observed percentage to meet the hydration goal than assignment to receive no reminder.” Do not automatically claim that the same effect would occur for every person or in every setting; that is a question about generalization, not the causal conclusion for this experiment.
Worked Example: An Experiment Without Random Assignment
Worked Example: An Experiment Without Random Assignment
A community center introduces a new stretching class at one location but not another. After six weeks, 36 of 60 participants at the location with the class report improved flexibility, compared with 20 of 50 participants at the other location. The director concludes that the class caused the improvement. Does the design justify that conclusion?
The explanatory variable is whether participants attended a location offering the new class; the response is whether they reported improved flexibility. The center introduced a program, so a treatment was imposed at one location. However, participants were not randomly assigned to the two locations or conditions. The groups may have differed before the class began.
The improvement percentage at the class location is \(36/60=0.60\), or 60%. At the other location it is \(20/50=0.40\), or 40%. The observed difference is 20 percentage points. The figures show an association between attending a location with the class and reporting improvement, but the comparison does not isolate the class as the cause.
For example, participants at the two locations could differ in their prior flexibility, motivation, schedules, or other relevant characteristics. These possibilities are not established facts about the participants; they illustrate why the lack of random assignment matters. The director can report the difference in these groups but should not claim that the class caused it based on this design alone.
A stronger experiment could randomly assign willing participants to the class or a comparison condition, then measure flexibility consistently in both groups. If the study is designed and carried out appropriately, random assignment would make a cause-and-effect conclusion more defensible.
Common Mistakes and AP Exam Tips
- Treating any difference between groups as proof of cause. A difference is a result to describe; the way the groups were formed determines whether a causal claim is supported.
- Calling a study experimental just because it follows people over time. A prospective study can still be observational. Ask whether researchers assigned treatments.
- Confusing random selection with random assignment. Random selection concerns choosing participants; random assignment concerns putting experimental units into treatment groups. Only the latter directly supports cause-and-effect reasoning.
- Assuming a large sample removes confounding. More observations do not make self-selected or otherwise different groups equivalent. Sample size does not replace random assignment.
- Writing “proves” when a study supports a causal conclusion. Even a randomized experiment is subject to chance variation and possible study limitations. Say that the results “support” or “provide evidence for” a cause-and-effect conclusion.
- Making a conclusion broader than the study. Name the treatment, response, and groups actually studied. Do not claim the effect applies to everyone unless the sampling design supports that separate inference.
A strong AP response states the design feature and links it to the conclusion. For an observational study, write that the variables are associated and explain that treatments were not randomly assigned, so cause and effect cannot be established. For a randomized experiment, identify the random assignment and state that the results support a cause-and-effect conclusion about the treatments compared.
Check Your Understanding
For each situation, identify the study design and state whether a cause-and-effect conclusion is justified.
- Researchers record how much time students choose to spend practicing an instrument and compare it with their performance scores. Can they conclude that more practice caused higher scores? Explain.
- A team randomly assigns volunteers to use either a new or standard plant-watering schedule, then compares plant growth. What feature of the design supports a cause-and-effect conclusion?
- A school randomly selects students to complete a survey about sleep and concentration. Does random selection make an association between sleep and concentration causal? Why or why not?
- A city introduces a new crosswalk design in one neighborhood and compares reported near-misses with a different neighborhood. Name one reason the observed difference may not establish that the design caused a change.
- What is the difference between random assignment and random selection, and what kind of conclusion does each help support?