Tutorials › AP Statistics › Why Correlation Does Not Imply Causation

Correlation · Tutorial 832 of 1000

Why Correlation Does Not Imply Causation

Distinguish describing an association from making a causal claim, and learn why study design matters when deciding whether cause and effect are supported.

Intermediate 9 min read

What You'll Learn

  • Explain why an observed association alone cannot establish cause and effect.
  • Analyze the ice cream sales and drownings example without claiming that one causes the other.
  • Distinguish observational studies from experiments.
  • Explain how random assignment can support a causal conclusion.
  • Separate causal conclusions about study subjects from generalizations to a wider population.

Association Is Not the Same as Cause and Effect

In “How Outliers Change the Correlation,” you saw how one observation can change the value of \(r\). Correlation summarizes the direction and strength of the linear association between two quantitative variables, but it does not explain why the variables are associated. Even a strong association does not, by itself, show that changes in one variable cause changes in the other.

This distinction matters whenever we read a claim such as “more of \(x\) leads to more of \(y\).” A scatterplot or correlation can show that larger values of \(x\) tend to occur with larger values of \(y\). That is a description of the data. A causal claim goes further: it says that changing \(x\), while other relevant circumstances are held appropriately comparable, produces a change in \(y\).

Definition: An association is a pattern of relationship between variables in observed data. A causal relationship means that changing one variable would produce a change in another variable. Association alone does not establish causation.

A causal claim is not just a stronger way to describe a scatterplot. It asks what would happen to the response if the explanatory variable were deliberately changed. Observational data usually show what happened among people, places, or other units as they were; they do not automatically tell us what would have happened to those same units under a different condition.

Ice Cream Sales and Drownings

Suppose a fictional town tracks weekly ice cream sales and the number of drownings at local swimming areas. In the town’s observations, weeks with higher ice cream sales also tend to have more drownings. These variables have a positive association. It would be a mistake to conclude that buying ice cream causes drownings.

The association may occur because both variables tend to be higher during warm, busy summer weeks. More people may buy cold treats, and more people may also go swimming. The town’s data show that the two measurements move together; they do not show that changing ice cream sales would change the number of drownings. This example illustrates why a relationship between two variables does not settle the cause-and-effect question.

The point is not that ice cream sales and drownings cannot be associated. They can be associated in the observed data. The point is that the association alone does not tell us which variable, if either, causes the other, or whether a different factor helps explain the pattern. Identifying and analyzing such additional variables is an important next step in understanding associations.

Worked Example: Ice Cream Sales and Drownings

Imagine a fictional coastal town that records weekly ice cream sales, in hundreds of dollars, and the number of drownings reported at nearby swimming areas. Over several summer weeks, the town notices that weeks with greater ice cream sales also tend to have more drownings.

Describe the association: In these observed weeks, ice cream sales and drownings have a positive association: higher sales tend to occur in the same weeks as more drownings. This statement describes a pattern in the paired data without claiming that the pattern is causal.

Evaluate the causal claim: The records do not show that increasing ice cream sales would increase drownings. Neither variable was assigned by the town, and the association alone cannot distinguish a direct causal effect from other explanations for the pattern. For example, hotter weeks could bring more ice cream purchases and more swimming activity.

Conclusion: The observations support describing a positive association in this fictional town during the weeks recorded. They do not establish that ice cream sales cause drownings. The responsible conclusion is about association, not cause and effect.

Observational Studies and Experiments

Study design helps determine what conclusions are justified. In an observational study, researchers observe or measure variables for the units in the study without assigning a treatment or deliberately imposing a condition. An observational study can reveal an association and can be useful for describing patterns. But by itself, it usually cannot establish that one variable caused changes in another.

In an experiment, researchers deliberately assign treatments or conditions to the units in the study and measure the response. If the experiment uses random assignment to decide which units receive which treatment, the groups can be compared in a way that supports a causal conclusion. Random assignment helps make the groups comparable, on average, with respect to factors that might affect the response.

Key distinction: Random assignment is about how units are placed into treatment groups. It can support a cause-and-effect conclusion. Random sampling is about how units are selected from a population. It supports generalizing results to that population. These are different steps and support different conclusions.

Random assignment does not mean every group will be identical in every respect. It makes the assignment process impartial, so pre-existing differences are not systematically used to place particular units in one group. With chance involved, some differences between groups can still occur. A careful causal conclusion describes the experiment and its conditions rather than promising that the treatment will have the same effect for every person or in every setting.

Random sampling and random assignment are often confused. A study can randomly assign volunteers to treatments without randomly selecting those volunteers from a larger population. That design may support a causal conclusion for the study conditions, but it does not automatically justify generalizing the results to everyone in the population. Conversely, a random sample in an observational study can help make a sample representative, but random selection alone does not establish causation.

Worked Example: Observing Study Habits and Exam Scores

A fictional school asks 120 students how many hours they studied for a biology exam and records each student’s exam score. Students who reported more study hours tended to have higher scores. The school did not assign students a number of study hours; each student reported their own study time.

Identify the design: This is an observational study because researchers recorded students’ existing study habits and scores rather than assigning study time. The data show an association between reported study hours and exam scores in these students.

Assess the causal claim: The association does not, on its own, prove that increasing study time would cause a higher score for a particular student. The students were not randomly assigned to study for different lengths of time, and the data do not establish why the pattern occurred. Students who choose to study longer may differ in other ways that are related to their scores.

Conclusion: It is appropriate to say that, in this fictional group, more reported study time was associated with higher exam scores. It is not justified to say that this observational study proves that additional study hours caused higher scores. A study designed to investigate a causal question would need an appropriate experiment, when one is practical and ethical.

How Random Assignment Supports a Causal Claim

Consider a fictional experiment about a new watering schedule for seedlings. A researcher has 60 similar seedlings and wants to compare a daily watering schedule with a less frequent schedule. The researcher randomly assigns 30 seedlings to each schedule, keeps other growing conditions the same as far as possible, and measures growth after four weeks.

Suppose the seedlings assigned to daily watering grow an average of 18.2 centimeters, while those assigned to the less frequent schedule grow an average of 15.1 centimeters. The observed difference in sample means is \(18.2-15.1=3.1\) centimeters. The difference describes these experimental groups; it is not by itself a guarantee that daily watering will increase growth by exactly 3.1 centimeters in every future group of seedlings.

Because the researcher deliberately assigned watering schedules at random, it is reasonable to investigate whether the watering treatment caused a difference in growth under these experimental conditions. The random assignment makes a causal comparison more credible than simply observing seedlings that happen to receive different amounts of water. Any conclusion should still name the seedlings and conditions studied, and should not automatically claim the same result for every plant species or growing environment.

Worked Example: A Randomized Seedling Experiment

In a fictional greenhouse experiment, 60 seedlings are randomly assigned to one of two watering schedules. Thirty receive a daily schedule, and 30 receive a less frequent schedule. All seedlings are grown for four weeks under the same lighting and temperature conditions. The daily-schedule group has a mean growth of 18.2 centimeters; the less-frequent group has a mean growth of 15.1 centimeters.

State the question: The researcher wants to know whether the assigned watering schedule affects seedling growth over four weeks in this greenhouse setup.

Plan by checking the design: This is an experiment because the researcher assigns the watering schedules. Random assignment is used to place seedlings into the two groups, so the treatment comparison can support a causal conclusion. The researcher also holds lighting and temperature conditions the same to make the comparison more focused. The study does not say that the seedlings were randomly sampled from all plants, so it does not establish that the result applies to every plant or growing environment.

Do the comparison: The observed difference in mean growth is \(18.2-15.1=3.1\) centimeters. The seedlings assigned to the daily schedule grew 3.1 centimeters more on average than those assigned to the less frequent schedule in this experiment.

Conclude in context: The randomized experiment provides a basis for concluding that the watering schedule caused a difference in mean seedling growth under the conditions studied, subject to the limitations of this particular experiment. It does not show that every daily-watered seedling grows more, nor does it establish the effect for other species or settings.

What Association Can and Cannot Tell You

An association can be valuable even when it does not establish cause and effect. It can help describe a population, identify a pattern worth investigating, or make predictions for cases similar to those observed. As in “Interpreting a Correlation in Context,” describe the variables and the observed linear pattern carefully. Do not add a causal explanation that the data and study design do not support.

When you encounter a causal statement, ask what the study actually did. Were variables simply recorded as they naturally occurred, or did researchers assign a treatment? If there was an experiment, were units randomly assigned? How were the units selected, and to whom might the findings apply? These questions separate evidence about association, evidence about causation, and evidence for generalization.

A causal experiment also needs to be ethical and practical. Researchers cannot randomly assign people to harmful exposures just to find out what happens. For many questions involving health, behavior, or community conditions, observational studies may be the available source of evidence. Such studies can show associations and inform further investigation, while their design limits how confidently a causal claim can be made.

Common Mistakes and AP Exam Tips

  • Writing “causes” when the study only observed variables. For an observational study, say that the variables are associated. A causal claim needs support from a design that assigns treatments, typically with random assignment.
  • Treating a strong correlation as proof of causation. The size of \(r\) describes the strength of a linear association, not the cause of that association. Even a correlation close to \(1\) or \(-1\) does not establish cause and effect.
  • Claiming that random sampling proves causation. Random sampling can support generalization from a sample to a population. Random assignment can support a causal comparison. Name which kind of randomization the study used.
  • Overstating what an experiment shows. A randomized experiment can support a causal conclusion for its study conditions, but it does not guarantee an identical effect for every unit or justify generalizing to populations that were not represented.
  • Explaining an association as if a possible explanation were proven. In the ice cream example, warmer weather and swimming activity are plausible explanations to consider. The observed association alone does not prove which explanation is responsible.

A strong AP response identifies the study design, states the conclusion that design supports, and limits the claim to the appropriate context. For example: “The study was observational, so it shows an association between the recorded variables but does not establish that one caused the other.” For a randomized experiment, name the treatment, the response, and the experimental conditions in the causal conclusion.

Key takeaway: Association describes how variables vary together; causation says that changing one variable produces a change in another. An association, even a strong one, does not establish causation. Observational studies show patterns, while experiments with random assignment can support causal conclusions under the conditions studied.

Check Your Understanding

For each situation, distinguish what the observed data show from what a causal claim would require.

  1. A fictional city finds that weeks with more outdoor concerts also have more emergency-room visits for heat illness. What can the city say about the association, and why does that alone not establish causation?
  2. A researcher records students’ sleep hours and quiz scores without assigning sleep schedules. Is this an observational study or an experiment? What kind of conclusion is supported?
  3. A greenhouse researcher randomly assigns plants to two light schedules and compares their growth. Which feature of the design supports a causal conclusion?
  4. Why does randomly selecting survey participants not, by itself, establish that a variable causes another variable?
  5. A randomized experiment uses volunteers rather than a random sample of the population. What can random assignment support, and what broader claim may remain unsupported?