Observing Patterns Without Assigning Treatments
In Parameters and Statistics, you learned to distinguish a population summary from a sample summary. When collecting data, another important question is how the groups or conditions being compared came about. Did investigators simply observe what people already did, or did they assign people to different treatments? That distinction affects what the study can support.
An observational study collects data about individuals or other observational units without deliberately assigning a treatment to change their responses. Investigators may ask people about their habits, measure conditions in their surroundings, or use records of events that have already occurred. The data can reveal a pattern or association between variables, but an observational study by itself does not establish that one variable caused a change in another.
An association exists when the values or categories of one variable tend to vary in a pattern with those of another. For example, people in one group might have a higher proportion of a particular outcome than people in another group. This describes what the data show. A causal claim goes further: it says that changing one variable would produce a change in the other.
The difference matters because groups in an observational study can differ in many ways besides the variable of interest. A third variable, sometimes called a lurking variable, may be related to both the explanatory variable and the response variable. If so, it can help explain the observed pattern. When the effects of variables are tangled together so that their separate contributions cannot be distinguished, the variables are confounded.
What an Observational Study Can Show
A study can still be useful even when it cannot establish cause. Observational data can show that two variables are associated in a particular group. They can help describe a population, identify patterns worth investigating, or suggest questions for future studies. A careful conclusion names the groups or variables and describes the observed relationship without saying that one variable made the other change.
For example, if students who use a study app tend to earn higher quiz scores, the data may show an association between app use and quiz scores. That does not prove that the app raised the scores. Students who choose to use the app may differ from students who do not: they might study more generally, have more time, or already feel confident in the subject. Those differences could also be related to quiz scores.
Knowing which event occurred first can be helpful, but it does not solve every problem. If an exposure happened before an outcome, that rules out the outcome as the cause of the earlier exposure. It does not rule out other explanations. For instance, a third factor may have influenced both the exposure and the later outcome.
A study may use a random sample and still be observational. Random selection is about how individuals enter the sample; it can help make the sample more representative of a defined population. It is different from assigning sampled individuals to treatments. A random sample does not, by itself, remove confounding or make an association causal.
A Practical Way to Describe an Observational Finding
When you read a study or write about its results, first identify how the data were collected. Then identify the variables and the pattern being compared. Finally, consider whether the study assigned treatments or merely observed existing conditions. This helps keep the conclusion within what the design supports.
Name the individuals or other observational units and the variables measured.
If investigators recorded existing conditions without assigning treatments, the study is observational.
Use the data to state how the outcome differs across groups or how the variables vary together.
Report the observed association in context, and do not claim that one variable caused the other. Consider plausible alternative explanations.
Worked Example: Outdoor Time and Sleep
Worked Example: Outdoor Time and Sleep
In a fictional survey, 80 teenagers report how much time they usually spend outdoors after school and whether they usually sleep at least eight hours on school nights. The investigators do not ask anyone to change their routine. Among 40 teenagers who report at least one hour outdoors each day, 28 usually sleep at least eight hours. Among 40 teenagers who report less than one hour outdoors each day, 18 usually sleep at least eight hours. Describe the finding and decide whether it supports a causal claim.
Identify the study design. The investigators recorded teenagers’ existing outdoor time and sleep habits. They did not assign teenagers to spend different amounts of time outdoors. This is an observational study.
Compare the sample proportions. The proportion reporting at least eight hours of sleep is \(28/40\) in the group with more outdoor time and \(18/40\) in the group with less outdoor time:
The difference is \(0.70-0.45=0.25\), or 25 percentage points. In this sample, 70% of teenagers who reported at least one hour outdoors usually slept at least eight hours, compared with 45% of those who reported less outdoor time.
State what the data support. The sample shows an association between reported outdoor time and reported sleep duration: the group reporting more outdoor time had a higher proportion who usually slept at least eight hours.
Limit the conclusion. The study does not show that spending more time outdoors caused teenagers to sleep longer. For example, school schedules, organized activities, or family routines could be related to both outdoor time and sleep. Those are possible alternative explanations, not facts established by these data.
Worked Example: Study-App Use and Quiz Scores
Worked Example: Study-App Use and Quiz Scores
For a fictional classroom investigation, a teacher records whether each of 60 students chose to use a study app during the week before a quiz. The teacher also records each student’s quiz score. The 30 students who chose to use the app had a mean score of 84 points; the 30 who did not use it had a mean of 78 points. The teacher did not assign app use. Describe the result and explain what it cannot establish.
Classify the study. Students chose whether to use the app, and the teacher recorded those choices and quiz scores. Since no treatment was assigned, this is an observational study.
Compare the sample means. Subtract the mean score for students who did not use the app from the mean for students who did:
In this sample, app users had a mean quiz score 6 points higher than nonusers. That is an association between app use and quiz score in the students observed.
Explain the limitation. It would not be justified to conclude that using the app raised quiz scores by 6 points. Students who chose the app may also have spent more time studying, had different prior knowledge, or been more motivated. These possible lurking variables could be related to both app use and quiz performance. The observed difference cannot separate the app’s possible effect from those other differences.
The fact that app use happened before the quiz is relevant to the sequence of events, but it does not remove the possibility of confounding. A careful conclusion describes the higher mean among app users without attributing that difference to the app.
Worked Example: Community Gardens and Fruit Intake
Worked Example: Community Gardens and Fruit Intake
In a fictional survey of 100 adults in a town, investigators record whether each person participates in a community garden and whether the person reports eating at least two servings of fruit each day. Of 50 garden participants, 35 report eating at least two servings. Of 50 nonparticipants, 20 report eating at least two servings. The investigators do not assign anyone to participate in a garden. Describe the association and give a reason it cannot establish cause.
Compare the proportions. Among garden participants, the proportion reporting at least two servings is \(35/50=0.70\). Among nonparticipants, it is \(20/50=0.40\). The difference is \(0.70-0.40=0.30\), or 30 percentage points.
Describe the pattern in context. In this sample, 70% of community-garden participants reported eating at least two servings of fruit per day, compared with 40% of nonparticipants. The data show an association between garden participation and reported fruit intake.
Consider another explanation. People who already place a high priority on healthy eating might be more likely to join a garden and to eat fruit regularly. Access to fresh food or other community activities might also be related to both variables. The survey did not assign garden participation, so it cannot establish that joining a garden caused higher fruit intake.
The possibility that people with a particular habit are more likely to enter a group is one reason to be cautious about causal interpretations. The group difference is real in the sample as described, but its cause is not determined by the observed association alone.
Common Mistakes and AP Exam Tips
- Turning “associated with” into “caused.” Saying that app use was associated with higher scores describes the observed pattern. Saying that app use raised scores makes a causal claim that this observational study does not support.
- Ignoring how groups formed. If people chose their own behavior or already belonged to a group, the groups may differ in other ways. Ask whether investigators assigned treatments or only recorded existing conditions.
- Assuming a random sample proves cause. Randomly selecting people for a sample is not the same as assigning them to treatments. Sampling affects who is represented; it does not, on its own, eliminate confounding.
- Calling every possible alternative explanation a proven confounder. When the study does not measure a third variable, describe it as a possible lurking variable or possible explanation. Do not claim the data demonstrate that it produced the association.
- Reporting a difference without naming the groups or outcome. A complete comparison identifies both groups, the variable being compared, and the relevant values or direction of the pattern.
- Thinking that a large difference proves causation. A striking difference may be important to describe, but its size does not change how the data were collected. Observational design still limits causal conclusions.
For a strong AP response, state the observed pattern with evidence and then connect the limitation to the study design. For example: “In this sample, app users had a mean quiz score 6 points higher than nonusers. Because students chose whether to use the app, the study is observational; differences such as study time or prior knowledge could help explain the association, so the results do not establish that the app caused higher scores.”
Check Your Understanding
For each situation, identify whether it is observational and describe what conclusion the data can support.
- A researcher records the amount of sleep and number of missed classes for a sample of students without asking them to change either. What kind of study is this, and what could an association between the variables show?
- A survey finds that people who own bicycles report more weekly exercise than people who do not. Give one possible lurking variable and explain why the survey does not prove that owning a bicycle causes more exercise.
- A town randomly selects households for a survey about home gardens and vegetable use. Does random selection by itself establish that gardening causes greater vegetable use? Explain.
- In a fictional sample, 24 of 40 people who use a meal-planning service report preparing dinner at home at least five nights a week; 14 of 40 nonusers report doing so. Calculate and compare the two sample proportions. State one conclusion the data do not justify.
- Why does knowing that an observed behavior occurred before an outcome help with causal reasoning but fail to rule out every alternative explanation?