Tutorials › AP Statistics › Common Errors Linking Regression to Causation

Regression and context · Tutorial 937 of 1000

Common Errors Linking Regression to Causation

Practice correcting regression statements that confuse association with causation, overstate what a model establishes, or claim more than a study design supports.

Intermediate 11 min read

What You'll Learn

  • Identify when a regression describes an observational association rather than a causal effect
  • Replace causal verbs with accurate associational wording when treatments were not assigned
  • Explain why a high r-squared does not show that one variable caused changes in another
  • Use random assignment to assess when a causal conclusion is supported
  • Limit causal and population claims to the cases and design that produced the data

When a Regression Statement Goes Too Far

A regression line can summarize a pattern and predict a response, but its equation does not tell you whether changing the explanatory variable would change the response. That depends on how the data were collected. In “Association Versus Causation in Regression” and “Experiments and Establishing Causation,” we distinguished observational studies from experiments. This tutorial uses that distinction to diagnose and repair common errors in graded-style examples.

The central task is to match the strength of a sentence to the evidence. If researchers merely record two variables, a regression can describe their association; it cannot, by itself, establish that one variable causes the other. If researchers use a well-designed randomized experiment to assign treatments, the design can support a cause-and-effect conclusion about the treatments. Even then, avoid claims about every individual or a wider population unless the evidence supports those claims.

Definition: A causal overclaim is a statement that says or implies that changing an explanatory variable causes a change in the response when the study design does not justify that conclusion, or when the statement reaches beyond the cases and outcomes the design can support.

A useful new technique is a causal-claim audit. Check four things before accepting or writing a regression conclusion: what the study did, what the model describes, what alternative explanations remain, and how narrowly the conclusion should be worded. This separates a correct numerical interpretation from an unsupported causal story.

1
Identify the design.
Were the variables observed, or did researchers deliberately assign a treatment using chance?
2
Read the model accurately.
State what the slope, prediction, or \(r^2\) describes, with the variables and units in context.
3
Check the causal wording.
For observational data, use language about association. For a well-designed randomized experiment, a causal conclusion about assigned treatments may be supported.
4
Set the boundary.
Keep the conclusion within the cases, treatment, response, and population the study can address.

This audit does not replace interpreting the model. As in “Keeping Context in Every Regression Sentence,” name the cases and variables, include the units, and say what the fitted line predicts. The audit adds a further question: does the design justify the verb you chose?

Observational Regression: Describe the Pattern, Not a Cause

In an observational study, researchers record the explanatory and response variables without assigning the explanatory variable as a treatment. A slope may describe the predicted response change associated with a one-unit increase in the explanatory variable. It does not show what would happen if a person, school, or community deliberately changed that variable.

A confounding variable or a lurking variable may help explain the association. Reverse causation may also be plausible: what is called the response might affect the explanatory variable, or an earlier version of the response might do so. As covered in “Identifying a Confounding Variable in a Scenario” and “Reverse Causation and Direction of Effect,” naming a plausible alternative does not prove it is responsible. It shows why a causal claim needs more than a fitted line.

Key wording: For an observational regression, say that the variables are associated, or that the model predicts a response difference across explanatory-variable values. Do not say that changing the explanatory variable causes the predicted change.

Be especially careful with statements about \(r^2\). As explained in “Interpreting \(r^2\) in Context,” it describes the fraction of variation in the response accounted for by the linear relationship in the data. It is not the percentage of the response caused by the explanatory variable, the percentage of cases for which a cause is established, or a measure of how likely the causal story is.

Graded Examples: Find the Error and Repair the Claim

Worked Example: Screen Time and Sleep in an Observational Study

A fictional survey records weekday recreational screen time and sleep for 180 high-school students. Researchers do not assign screen time. Let \(x\) be a student's recreational screen time in hours per day and \(y\) be weekday sleep in hours per night. The fitted line is \(\hat{y}=8.4-0.6x\), and \(r^2=0.36\).

A student writes: “Every extra hour of screen time causes students to lose 0.6 hours of sleep. Screen time causes 36% of their sleep loss.” Grade and correct this statement.

Audit the design. This is an observational survey: screen time was recorded, not randomly assigned. The regression describes an association between screen time and sleep among the surveyed students. The design does not establish that reducing an individual student's screen time would increase that student's sleep by a specified amount.

Check the slope. The slope is \(-0.6\) hours of sleep per additional hour of screen time per day. Its sign indicates that higher recorded screen time is associated with a lower predicted amount of sleep. “Every extra hour ... causes” turns that association into a causal claim. “Students lose” also sounds like a change measured in each student, which the survey did not establish.

Check \(r^2\). The calculation is \(0.36 \times 100\%=36\%\). But that is a percentage of variation in sleep accounted for by the linear relationship with screen time in these data—not a percentage of sleep loss caused by screen time. The response variation has not been divided into causal and noncausal portions.

Full-credit repair: Among the 180 students surveyed, each additional hour of weekday recreational screen time per day is associated with a predicted decrease of 0.6 hours in weekday sleep per night. About 36% of the variation in weekday sleep among these students is accounted for by its linear relationship with recreational screen time. Because the data are observational, they do not establish that screen time causes less sleep.

This repair reports both model summaries accurately and explains the design limit. Plausible confounders might include school start time, workload, or a student's schedule; the survey alone cannot separate their possible contributions. The point is not to assert that any particular factor explains the pattern, but to avoid claiming that the regression has ruled out alternatives.

Worked Example: Exercise and Mood Scores

A fictional community center asks 95 adult participants to report their usual weekly exercise minutes and complete a mood questionnaire on the same day. Let \(x\) be usual exercise in minutes per week and \(y\) be the questionnaire score, with higher scores indicating more positive mood. The fitted line is \(\hat{y}=42+0.03x\).

A report says: “Increasing exercise by 100 minutes a week raises a person's mood score by 3 points. The regression proves that exercise improves mood.” Identify the correct arithmetic, the causal error, and a careful replacement.

Check the model statement. A 100-minute difference in \(x\) corresponds to a predicted score difference of \(0.03(100)=3\) points along the fitted line. The units check: the slope is questionnaire points per exercise minute per week, so multiplying by 100 minutes per week gives 3 questionnaire points. That arithmetic describes a predicted difference, not a measured change after an intervention.

Audit the design and timing. The center recorded both variables at one time; it did not assign exercise. The data therefore show an association, not that increasing exercise causes an individual's mood score to rise. Direction is also uncertain: a person's mood could affect how much they exercise. Other factors, such as health or available free time, could be related to both variables. These are plausible alternatives, not conclusions established by this survey.

Full-credit repair: Among the adult participants surveyed, people reporting 100 more minutes of usual weekly exercise are predicted by the fitted line to have mood scores 3 points higher, on average. Because exercise was not assigned and both variables were recorded at the same time, this association does not establish that increasing exercise causes mood scores to improve.

The phrase “people reporting ... are predicted ... to have” keeps the statement tied to the observed comparison. If the question is specifically about an intervention, the survey is not enough to answer it. A randomized experiment could provide stronger evidence about the effect of an assigned exercise program, although its conclusion would still depend on the experiment's participants and implementation.

Worked Example: A Randomized Workshop and Test Scores

A fictional school recruits 120 students who volunteer for a study of a study-skills workshop. Researchers randomly assign 60 students to the workshop and 60 to usual advisory. Both groups take the same posttest under the same conditions. The control group's mean score is 68 points and the workshop group's mean is 74 points. Code \(x=0\) for usual advisory and \(x=1\) for assignment to the workshop; let \(y\) be posttest score in points. For this binary \(x\), the fitted line is \(\hat{y}=68+6x\), since the slope is \(74-68=6\) points.

A student concludes: “The workshop caused every student to score 6 points higher, and this proves the workshop raises scores for all students.” The first part recognizes that the experiment can support causal reasoning, but the full statement makes two errors.

Audit the design. Chance assigned the workshop treatment, and the study used the same posttest conditions for both groups. Assuming the assignment and implementation were carried out as described, the comparison supports a causal conclusion about assignment to the workshop for the students in this experiment. The regression slope is the difference between the two group mean scores: \(74-68=6\) points. Random assignment helps make the groups comparable, on average, before treatment.

Correct the individual-level claim. The 6-point difference is a difference in group means, not a promise that every workshop student gained 6 points. The study does not compare each student's score under both conditions; an individual student's outcome under the treatment they did not receive is unknown. The result supports a conclusion about the average effect of assignment in the experiment, not an identical effect on every student.

Correct the population claim. The 120 students volunteered; they were not randomly selected from all students. Random assignment supports the cause-and-effect comparison for the experimental groups, but it does not by itself make the volunteers representative of every student. As discussed in “Population Scope and Generalizing Results,” random assignment and random selection answer different questions.

Full-credit repair: For the 120 students in this experiment, assignment to the study-skills workshop produced a group mean posttest score 6 points higher than assignment to usual advisory. Because students were randomly assigned and testing conditions were the same, the experiment supports a causal conclusion about the average effect of workshop assignment for these participants. Since the participants volunteered, the result cannot automatically be generalized to all students, and it does not mean every student gained 6 points.

Notice the boundary: the causal claim is about assignment to the workshop, the measured posttest score, and the participants studied. If some students did not attend the sessions they were assigned, the experiment still directly compares assignment groups; it would not automatically establish the effect of actually attending for every student. Match the claim to what researchers controlled.

Common Mistakes and What Full Credit Says

A grading response should do more than write “correlation is not causation.” It should connect the limitation to the specific design and replace the flawed sentence with a defensible statement.

  • Using causal verbs for recorded variables. “Causes,” “leads to,” “makes,” and “increases” can imply an intervention. For an observational study, identify the association and state that the study does not establish causation.
  • Calling the slope an individual change. A regression slope describes a predicted response difference per explanatory-variable unit. It does not show that one particular person would change by exactly that amount if their \(x\) value changed.
  • Turning \(r^2\) into a causal percentage. Say what percentage of response variation is accounted for by the linear relationship. Do not call it the percentage caused, explained by a mechanism, or correctly predicted.
  • Assuming random assignment means random sampling. Random assignment supports comparing treatments for causal evidence. Random selection is what can support generalizing to a population, if the sampling process and population frame are appropriate.
  • Overstating experimental results. A randomized experiment can support a causal conclusion about the assigned treatment and the study outcome. It does not establish that every individual responds identically, that a different version of the treatment has the same effect, or that results apply to people not represented in the study.
  • Listing a confounder as if it were proven. Say that a third variable could help explain an observational association and explain how it may relate to both variables. Do not claim the study showed that this variable caused the pattern unless it measured and analyzed that question appropriately.
  • Leaving out the context when correcting the claim. Name the cases, explanatory variable, response, units, and relevant design feature. A correction that is only “there may be confounding” does not interpret the regression or answer what the slope means.

A strong AP response often follows this pattern: “In this observational study, [explanatory variable] is associated with a predicted [increase or decrease] in [response], in [units] and in context. Because researchers did not randomly assign [explanatory variable or treatment], the data do not establish that changing it causes the response to change.” For a randomized experiment, name the assigned treatment, response, comparison, and the group or population the conclusion can address. Avoid adding claims about individual effects or broader populations without evidence.

Key takeaway: A regression describes a relationship; the study design determines whether a causal conclusion is justified. For observational data, repair causal language into an in-context association. For a well-designed randomized experiment, make a bounded causal claim about the assigned treatment and measured response—not a guarantee for every individual or proof about a broader population.

Check Your Understanding

For each item, identify the design issue and write a corrected statement that stays within the evidence.

  1. A survey records weekly gaming hours and hours of sleep for teens. A student says, “Each extra gaming hour causes 0.4 fewer hours of sleep.” What wording is supported, and why?
  2. A regression of commute time on distance has \(r^2=0.64\). Explain why “distance causes 64% of commute time” is not an appropriate interpretation.
  3. A city randomly assigns participating households to one of two reminder-message schedules and compares recycling rates. What feature supports a causal conclusion about the schedules? What does random assignment alone not establish?
  4. In a randomized experiment, the treatment group's average score is 5 points higher than the control group's average. Why is “every treated person scored 5 points higher” not justified?
  5. Name one additional variable that could be related to both exercise and mood in an observational study. Explain why identifying it as a possibility is not the same as proving it caused the observed association.