From a Regression Result to a Headline
A news headline often compresses a study into a few words. That compression can turn “these variables were associated” into “one thing caused another,” or make a result from a narrow sample sound true for everyone. A regression equation may help describe the data, but the study’s design determines what conclusions the evidence can support.
In “Common Errors Linking Regression to Causation,” we practiced identifying causal overclaims. Here, we use a more focused technique: a headline audit. The goal is not to decide whether a headline sounds plausible. It is to compare the headline’s precise claim with the variables measured, the association observed, and how the cases were selected and studied.
The examples below use fictional news-style headlines and invented study results. They are not reports of real studies. In each case, ask three connected questions: What association does the regression summarize? Did researchers observe the variables or assign a treatment? Who were the cases, and how were they selected?
Identify the cases, explanatory variable, response variable, units, and the comparison the headline makes.
Read the slope and, when available, \(r\) or \(r^2\) in context. Make sure the headline does not turn a predicted difference into a guaranteed individual change.
Ask whether the study recorded variables or used random assignment. Consider plausible alternative explanations when the data are observational.
Check which cases were studied and whether the selection method supports extending the conclusion to a broader population.
A headline’s verbs are clues, not proof. “Associated with” usually makes an associational claim. Words such as “causes,” “cuts,” “boosts,” and “prevents” make causal claims. “Linked to” can be ambiguous, so read the article’s supporting sentence and study description rather than judging the headline in isolation.
Worked Headline Audits
Worked Example: Street-Tree Shade and Pavement Temperature
Fictional headline: “Planting More Street Trees Cools City Blocks.” A fictional report describes a study of 24 downtown blocks. Researchers recorded each block’s existing tree shade and measured pavement temperature at the same time of day. The blocks were chosen because they were accessible, not randomly selected, and researchers did not plant trees or assign shade. Let \(x\) be shade coverage in percentage points and \(y\) be pavement temperature in degrees Celsius. The fitted line is \(\hat{y}=38.6-0.18x\), with \(r=-0.62\).
Check the association. The slope is \(-0.18\) degrees Celsius per percentage point of shade. For a 10-percentage-point difference in shade, the model predicts a temperature difference of
Thus, among the observed blocks, the fitted line predicts a pavement temperature 1.8 degrees Celsius lower for blocks with 10 percentage points more shade. This is a comparison described by the model, not a measured cooling effect from planting trees.
The coefficient of determination is
About \(38.44\%\), or \(38.4\%\) rounded, of the variation in pavement temperature among these blocks is accounted for by its linear relationship with shade coverage in these data. As covered in “Interpreting \(r^2\) in Context,” that percentage is not the amount of cooling caused by shade.
Check the design and scope. This is an observational study: existing shade and temperature were recorded, and no treatment was assigned. A block’s surface materials or traffic could be related to both shade and temperature; these are plausible alternative explanations, not proven causes of the pattern. Also, the blocks were chosen for accessibility, so the results do not automatically describe every city block. “Planting” makes the headline sound like an intervention whose effect was tested, but it was not.
Defensible rewrite: Among the 24 accessible downtown blocks studied, greater existing tree shade was associated with lower pavement temperatures. The fitted line predicts a 1.8-degree-Celsius lower temperature for a 10-percentage-point increase in shade coverage, but the observational study does not establish that planting trees causes the decrease.
Worked Example: Phone Notifications and a Focus Task
Fictional headline: “Phone Notifications Cut Workers’ Focus.” Researchers in a fictional study use app records and a short focus task for 160 adult volunteers who use a particular phone app. They record each person’s typical daily notification count and focus-task score. Nobody is assigned a notification level. Define \(x\) as the number of daily notifications measured in tens, and \(y\) as the focus-task score in points. The fitted line is \(\hat{y}=82-1.5x\), with \(r=-0.50\).
Translate the predictor scale. Since \(x\) is measured in tens of notifications, a change from 20 to 50 notifications per day is a change from \(x=2\) to \(x=5\). The predicted score difference is
The fitted line predicts a focus-task score 4.5 points lower for a person with 50 rather than 20 daily notifications. The slope does not say that reducing one person’s notifications will raise that person’s score by exactly 4.5 points. It summarizes an association across the observed cases.
Here,
So \(25\%\) of the variation in focus-task scores among these volunteers is accounted for by the linear relationship with their recorded notification counts. The remaining variation is not a measure of the percentage of people unaffected by notifications; \(r^2\) describes variation, not the share of individuals for whom a claim is true.
Check the design and scope. Because the researchers recorded rather than assigned notification counts, this study cannot establish that notifications caused lower scores. Workload or sleep, for example, could be related to both notification counts and focus-task scores. The direction could also be complicated: people having trouble concentrating might check their phones more often. These are possibilities to consider, not findings demonstrated by the regression. Finally, the volunteers all used one app, so the results should not automatically be generalized to all workers or phone users.
Defensible rewrite: Among the 160 adult app users who volunteered, higher daily notification counts were associated with lower focus-task scores. The fitted line predicts a 4.5-point difference between 20 and 50 notifications per day, but the observational study does not show that notifications cause an individual’s score to fall.
Worked Example: A Randomized Watering Reminder
Fictional headline: “Text Reminder Cuts Household Watering by 8 Liters a Week.” In a fictional experiment, 80 volunteer households are randomly assigned to receive either a weekly watering reminder text or standard gardening information. Researchers measure each household’s water use for outdoor watering during the following week, using the same method in both groups. The mean is 46 liters for the standard-information group and 38 liters for the reminder group. Let \(x=0\) represent standard information and \(x=1\) represent the reminder; let \(y\) be weekly outdoor water use in liters.
Check the numerical comparison. With a binary explanatory variable coded 0 and 1, the fitted line’s intercept is the control-group mean and its slope is the difference between the two group means. Here the slope is
The fitted line is therefore \(\hat{y}=46-8x\). It predicts 46 liters for the standard-information group and 38 liters for the reminder group. The 8-liter difference is a difference in group means, not a guarantee that every household will reduce its use by exactly 8 liters.
Check the design. Unlike the first two examples, this is an experiment: researchers randomly assigned households to a treatment. If assignment and implementation were carried out as described and the groups were measured comparably, the experiment supports a causal conclusion about the average effect of assignment to the reminder, for these participants and this measured week. The conclusion concerns assignment to receive the text, not necessarily the effect of reading or following it.
Check the scope. The households volunteered; they were not randomly selected from all households. As explained in “Population Scope and Generalizing Results,” random assignment supports causal comparison, while random selection is the feature that can support generalization to a population. The headline’s definite “cuts” can also sound like every household will save that amount.
Defensible rewrite: In a randomized experiment with 80 volunteer households, assignment to a weekly watering reminder produced a group mean outdoor water use 8 liters per household lower than assignment to standard gardening information during the measured week. The result does not guarantee an 8-liter reduction for each household or establish the effect for all households.
Common Mistakes and Full-Credit Wording
A strong critique does not stop at “correlation is not causation.” It identifies exactly what the model says, what the study did, and which part of the headline goes beyond the evidence. When the regression includes a slope or \(r^2\), interpret that statistic accurately before explaining the design limit.
- Treating a predicted difference as an individual change. A slope describes a predicted response difference for a change in the explanatory variable. It does not show that changing one individual’s \(x\) will produce that exact change in \(y\).
- Using causal verbs after an observational study. If researchers only recorded the variables, replace “causes,” “cuts,” or “boosts” with in-context language such as “was associated with” or “the model predicts a difference.” Then state that the study does not establish causation.
- Reading a scaled predictor incorrectly. If \(x\) is measured in tens, hundreds, or thousands, use that scale when calculating a predicted difference. State the units so readers can see what comparison the slope represents.
- Calling \(r^2\) a causal percentage. Say what percentage of the variation in the named response is accounted for by its linear relationship with the explanatory variable in the observed data. Do not describe it as the percentage caused or the percentage of cases affected.
- Assuming random assignment makes a sample representative. Random assignment and random selection serve different purposes. Assignment can support a causal comparison; it does not by itself justify generalizing to a wider population.
- Listing a possible confounder as a proven explanation. Name a plausible third variable and explain how it could be related to both variables. Do not claim that the study showed this variable caused the association unless the evidence supports that conclusion.
- Ignoring what the headline is actually claiming. A study of existing conditions does not automatically test an intervention such as planting, reducing, or increasing something. Check whether the study manipulated the explanatory variable or merely recorded it.
For a clear AP response, connect the numerical result to the cases and units, describe the design, and give the limit in the same context. For example: “Among the [cases] studied, [explanatory variable] was associated with [response]. The fitted line predicts [contextual difference]. Because researchers recorded rather than randomly assigned [explanatory variable], the study does not establish that changing it causes [response] to change.” For a randomized experiment, state the treatment comparison and average response, then keep the causal conclusion within the participants and conditions studied unless random selection supports a broader claim.
Check Your Understanding
For each fictional headline, identify the key association or design question and suggest wording that stays within the evidence.
- A survey records weekly gardening time and vegetable harvest for 90 volunteers. A headline says, “More Gardening Produces Bigger Harvests.” What can the regression describe, and what does the survey not establish?
- A fitted line predicts 2.4 more points on a quiz for every additional hour of practice. Explain why this is not a guarantee that each student who practices one more hour will gain 2.4 points.
- A report says \(r^2=0.49\) for the association between distance walked and daily step count. What does \(r^2\) describe, and what would be wrong with calling it “49% of steps caused by walking distance”?
- A randomized trial assigns volunteer households to one of two composting instructions. What does random assignment support, and why does it not automatically show the result applies to every household?
- A headline says “More Street Shade Lowers Temperatures,” based on observations of existing blocks. Name one plausible variable that could relate to both shade and temperature, and explain how it affects the headline’s causal claim.