A Fitted Line Describes an Association
A regression line summarizes how two variables vary together in the observed data. It can describe a pattern and provide predictions, but the line by itself does not explain why the pattern exists. In particular, a line does not prove that changing the explanatory variable would cause the response variable to change.
In “Describing the Context of a Regression Study,” you identified the cases, measurements, and roles of the variables. Here we use that context to make a careful distinction: a regression model may show an association between an explanatory variable and a response, while a claim of causation says that changes in one variable produce changes in the other. The second claim needs support beyond a fitted line.
For a least-squares regression line, the slope describes the change in the predicted response for a one-unit increase in the explanatory variable. That is a statement about the model’s pattern. It is not automatically a statement about what would happen if someone deliberately changed the explanatory variable.
This distinction matters even when the relationship is strong, the points lie close to the line, or the coefficient of determination is large. As discussed in “Interpreting \(r^2\) in Context,” \(r^2\) describes the fraction of variation in the response accounted for by its linear relationship with the explanatory variable. A high \(r^2\) does not identify the reason for that relationship.
Electricity Use and Outdoor Temperature
Suppose a building manager records the average outdoor temperature and electricity use for each of four weeks. The building uses electric heating. The following invented data are a small illustration, not results from a real study.
| Week | Average outdoor temperature, \(x\) (°C) | Electricity use, \(y\) (kWh per day) |
|---|---|---|
| 1 | 0 | 920 |
| 2 | 10 | 780 |
| 3 | 20 | 680 |
| 4 | 30 | 620 |
For these data, the least-squares line is \(\hat{y}=900-10x\). Here, \(\hat{y}\) is the predicted daily electricity use in kilowatt-hours, and \(x\) is the average outdoor temperature in degrees Celsius. The slope is negative: in these observed weeks, higher temperatures are associated with lower electricity use.
We can check the fitted line from the data. The mean temperature is \(15\) °C, and the mean electricity use is \(750\) kWh per day. The sum of the products of the deviations is \(-5000\), and the sum of squared temperature deviations is \(500\). Thus the slope is \(-5000/500=-10\), and the intercept is \(750-(-10)(15)=900\). The resulting line is \(\hat{y}=900-10x\).
The slope’s careful interpretation is: for each increase of 1 °C in average outdoor temperature, the fitted line predicts a decrease of 10 kWh in daily electricity use, on average, for weeks like those observed. This describes the association in these data. It does not establish that raising the outdoor temperature by 1 °C would cause electricity use to decrease by exactly 10 kWh per day.
A causal explanation is plausible: when it is warmer, a building with electric heating may need less heat. But the four recorded weeks do not isolate that effect. Week of the year, daylight, building occupancy, thermostat settings, and other electricity uses could also vary with temperature. The regression line does not separate those influences from one another.
What Would Make a Causal Claim More Credible?
The way data are collected matters. In an observational study, researchers record variables as they occur, without assigning the explanatory variable to the cases. A regression line from such data can summarize an association, but differences between cases or circumstances may help produce that association. A variable related to both the explanatory and response variables can be a possible alternative explanation.
In a randomized experiment, researchers deliberately assign treatments to cases using random assignment. If the experiment is well designed and carried out, this helps make the treatment groups comparable in other respects. A difference in outcomes can then provide evidence that the assigned treatment caused a difference in the response for the cases studied. Random assignment does not make every possible source of error disappear, but it strengthens the basis for a causal conclusion.
Random assignment and random sampling have different jobs. Random assignment helps support a cause-and-effect conclusion about the cases in an experiment. Random sampling can help support generalizing results to a larger population. One does not automatically provide the benefit of the other: randomly assigning volunteers to treatments does not, by itself, make those volunteers a representative sample of everyone.
Was it simply observed, or did researchers deliberately assign different values or treatments?
In an experiment, random assignment can help balance other influences across treatment groups and support a causal interpretation.
For an observational regression, describe the association. For a well-designed randomized experiment, a causal conclusion may be supported, with its strength depending on the evidence.
Consider how cases were selected when deciding whether results can extend beyond the cases studied.
Worked Examples: Association, Cause, and Study Design
Worked Example: Interpreting the Electricity Regression
Use the four-week electricity data above. A building manager says, “The regression line proves that warming the building’s outdoor environment by 1 °C will reduce its electricity use by 10 kWh per day.” Evaluate the statement and give a more appropriate interpretation.
Identify what the line says. The fitted line is \(\hat{y}=900-10x\). Its slope is \(-10\) kWh per day per degree Celsius. For each 1 °C increase in average outdoor temperature, the line predicts a 10 kWh-per-day decrease in electricity use, on average, among weeks like those observed.
Check the claim. The data record weeks as they occurred; the manager did not randomly assign temperatures or otherwise isolate their effect. Other conditions—such as occupancy, daylight, and heating settings—may differ across weeks. The fitted line therefore demonstrates an association in these observations, not proof that changing temperature causes the stated reduction.
Give a defensible conclusion. “In the four observed weeks, higher average outdoor temperatures were associated with lower daily electricity use. The fitted line predicts a decrease of 10 kWh per day for each 1 °C increase in temperature, but these observational data alone do not establish that temperature caused the decrease.”
This conclusion preserves the useful slope interpretation while avoiding a claim about what would happen under a controlled change in temperature.
Worked Example: Tutoring Hours and Exam Scores
A fictional school reviews records for students who chose how many hours of optional tutoring to attend. The fitted line relating tutoring hours \(x\) to exam score \(y\), in points, is \(\hat{y}=58+3.2x\). One student says, “An extra hour of tutoring causes a 3.2-point increase.” What does the line support, and what does it not establish?
Interpret the slope as an association. In these records, each additional tutoring hour is associated with a predicted exam score that is 3.2 points higher, on average. For a student with 5 tutoring hours, the line predicts \(58+3.2(5)=58+16=74\) points.
Consider how tutoring was chosen. Students selected their own tutoring hours; the school did not randomly assign them. Students who were already struggling might seek more tutoring, while students who are especially motivated might also attend more. Those differences could be related to scores, so the observed association does not by itself show the effect of adding an hour of tutoring.
State the conclusion carefully. “Among the students in these records, more tutoring hours were associated with higher predicted exam scores. The regression line does not prove that assigning a student one additional tutoring hour would raise that student’s score by 3.2 points.”
The prediction for 5 hours is a model estimate, not a promise about an individual student. The slope summarizes a pattern across the observed students; it does not specify what caused that pattern.
Worked Example: Random Assignment and Seedling Growth
In a fictional classroom experiment, eight seedlings of the same variety are randomly assigned: four receive standard lighting and four receive supplemental lighting. After the same growing period, their growth is measured in centimeters. The standard-light seedlings grow 8, 9, 10, and 9 cm; the supplemental-light seedlings grow 11, 12, 10, and 11 cm. What do the observed results show, and why is a causal interpretation more reasonable here?
Calculate the group means. The standard-light mean is \((8+9+10+9)/4=36/4=9\) cm. The supplemental-light mean is \((11+12+10+11)/4=44/4=11\) cm. The observed difference in means is \(11-9=2\) cm.
Connect the result to the design. Lighting was deliberately assigned, and random assignment helps make the two groups comparable in other respects. If the experiment was conducted consistently, the 2 cm difference in these group means provides evidence that supplemental lighting increased growth for these seedlings.
Keep the conclusion within the evidence. The observed difference does not mean that every seedling grows exactly 2 cm more under supplemental lighting. There are only four seedlings in each group, so the result is based on a small experiment. The design supports a causal interpretation more than the observational tutoring records do, but the size of the experiment limits how firmly we should state the conclusion and how broadly we should generalize it.
The important contrast is not whether a regression line was calculated. It is that the lighting treatment was assigned at random, allowing a more credible comparison of outcomes under the two conditions.
Common Mistakes and AP Exam Tip
A full-credit response states what the data and design warrant, rather than turning a model’s prediction into a causal promise.
- Using “causes” as a synonym for “is associated with.” For observational data, say that higher values of one variable tend to occur with higher or lower values of the other. Do not say that changing one variable produces the change.
- Reading the slope as a guaranteed individual effect. A slope describes a predicted average change in the response per unit increase in the explanatory variable. It does not guarantee that each individual’s response changes by that amount.
- Treating a strong fit as proof of cause. A high \(r^2\), a strong correlation, or small residuals describes how well the line summarizes the observed pattern. None of these identifies why the variables are related.
- Assuming a plausible explanation is proven. Electric heating offers a sensible reason for temperature and electricity use to be related. A plausible mechanism is not the same as evidence that rules out other explanations.
- Confusing random sampling with random assignment. Random sampling concerns how cases are selected; random assignment concerns how treatments are allocated in an experiment. Explain which one the study used and what it can support.
- Overstating experimental results. Random assignment can support a causal conclusion, but the conclusion should still reflect the observed outcomes, the experiment’s limitations, and the cases studied.
A reliable AP response first identifies whether the data are observational or experimental. Then it describes the association in context and explains whether the design supports a causal conclusion. For the electricity example, a full-credit answer says that warmer observed weeks were associated with lower electricity use, while noting that the regression line alone does not prove temperature caused the decrease.
Check Your Understanding
For each question, distinguish what an observed regression pattern shows from what the study design can establish.
- In the electricity example, what does the slope of \(-10\) kWh per day per degree Celsius predict? Why does it not prove that raising outdoor temperature causes the predicted decrease?
- A company finds a positive association between time spent using a study app and test scores in its user records. What is an appropriate association statement, and why would “the app caused higher scores” need more support?
- Why does a high \(r^2\) for an observational regression not establish a causal relationship?
- In the seedling experiment, what role did random assignment play? Would randomly selecting seedlings from a large population be the same thing?
- A well-designed randomized experiment finds higher average growth for a treatment group. What kind of conclusion can the design support, and what should still be considered before generalizing the result?