Read the Argument, Not Just the Numbers
A regression argument can contain correct output and still communicate a misleading conclusion. A writer might leave out what was measured, use the wrong meaning for \(r^2\), or claim that one variable causes another when the data only show an association. Reviewing an argument means checking how its words connect to the evidence—not merely checking whether its numbers look plausible.
In “Turning Output Into a Contextual Interpretation,” you practiced interpreting the slope, intercept, and \(r^2\). Here, use those interpretations as a standard for evaluating someone else’s writing. You can also draw on “Selecting Evidence for a Written Conclusion” to check whether each piece of evidence actually supports the claim being made, and on “Describing Limitations of a Regression Analysis” to make any criticism specific.
A useful new technique is to audit the argument in four passes. First, recover the context: who or what were the cases, and what were \(x\) and \(y\), with units? Second, check the statistical wording: does each number mean what the writer says it means? Third, match the evidence to the claim: does the statistic or graph directly address that claim? Fourth, check the limits: is the claim appropriate for the observed cases, \(x\)-values, and data-collection method?
Identify the cases, explanatory variable, response variable, units, setting, and any relevant observed \(x\)-range.
Ask whether slope, intercept, \(r\), \(r^2\), \(s\), and predictions are described according to their meanings.
Decide whether the evidence supports the exact statement, rather than a nearby or stronger statement.
Check the model’s scope and whether the data-collection method supports the writer’s conclusion, especially a causal one.
Three Questions That Expose Common Problems
A reviewer should be able to point to the exact phrase that needs attention. “This analysis is bad” is not a useful critique. Instead, identify what is missing or wrong and state why it matters. For example, a slope sentence without units is incomplete; describing \(r^2\) as the percentage of cases predicted correctly is incorrect; and saying an explanatory variable “caused” a change may go beyond what an observational regression can establish.
Context is more than naming two variables. A reader may also need to know which cases were studied, where or when the data were collected, and how the observations were obtained. These details help set the scope of the conclusion. As explained in “Scope of Inference for a Regression Model,” a model’s scope concerns both the cases and setting represented and the observed range of the explanatory variable.
Worked Example: Find Overclaims in a Commute-Time Analysis
Situation. In an invented observational project, students recorded commute time and the number of days each student arrived late during one school month for 36 students at a single school. Let \(x\) be commute time in minutes and \(y\) be late-arrival days per month. The observed commute times ranged from 5 to 42 minutes. The fitted line was \(\hat{y}=0.18+0.07x\), with \(r=0.74\), \(r^2=0.548\), and \(s=1.6\) late-arrival days.
A classmate writes: “Every extra minute of commuting causes 0.07 more late days. The model predicts 55% of students correctly. A student with a 60-minute commute should be late 4.38 days per month, so the school should use this model for students with long commutes.”
State. The argument includes an incorrect interpretation of \(r^2\), an unsupported causal claim, and a prediction outside the observed commute-time range. It also does not identify the cases and setting in its conclusion.
Plan. Check each statement against the output and study description. Interpret the slope in context without claiming cause, state what \(r^2\) measures, and compare the 60-minute value with the observed range.
Do. The slope is \(0.07\) late-arrival days per month for each additional minute of commute time. A suitable interpretation is: “Among these students, each additional commute minute is associated with a predicted increase of \(0.07\) late-arrival days per month according to the fitted line.” Because the project was observational, the slope does not show that longer commutes cause students to arrive late.
The reported \(r^2\) is \(0.548\), or \(0.548\times100\%=54.8\%\), rounded to one decimal place. It means that about 54.8% of the variation in late-arrival days among these students is accounted for by the linear regression of late-arrival days on commute time. It is not the percentage of students predicted correctly.
The calculation for the fitted value at 60 minutes is \(0.18+0.07(60)=4.38\) late-arrival days. However, 60 minutes is greater than the observed maximum of 42 minutes, so using the model there is extrapolation. The value is a model prediction, not a guarantee; \(s=1.6\) late-arrival days describes the typical size of residuals around the fitted line in this setting.
Conclude. A more defensible summary is: “For the 36 students at this school, commute time had a positive linear association with late-arrival days during the school month. The line accounts for about 54.8% of the variation in late-arrival days among these students. These observational data do not establish that commute time causes lateness, and the fitted value for a 60-minute commute is an extrapolation beyond the observed range of 5 to 42 minutes.”
Separate Missing Context from Incorrect Meaning
Not every weakness is a false statement. Some statements are mathematically correct but too incomplete to be useful. “The slope is 2.5” might accurately copy output, but it does not explain which response is predicted to change, per unit of which explanatory variable, or in what units. A reviewer should distinguish an incomplete interpretation from an incorrect one and describe the repair needed.
The intercept is another place where a statement can be mathematically accurate but practically misleading. It represents the model’s predicted response at \(x=0\). As in the earlier tutorial on interpreting regression output, check whether zero is in the observed range before treating that prediction as meaningful in context. The line has an intercept whether or not the data support a practical interpretation at zero.
Worked Example: Repair a Fertilizer Claim
Situation. In an invented greenhouse project, students observed 30 basil plants. They recorded fertilizer applied in tablespoons per week (\(x\)) and plant growth in centimeters per month (\(y\)). The students who tended the plants chose the fertilizer amounts; the amounts were not randomly assigned. Observed amounts ranged from 0.5 to 3.0 tablespoons per week. The output gives \(\hat{y}=1.2+2.5x\), \(r^2=0.64\), and \(s=1.1\) centimeters per month.
A classmate writes: “Fertilizer causes plants to grow 2.5 centimeters more each month per tablespoon. Fertilizer explains 64% of growth. With no fertilizer, plants grow 1.2 centimeters per month, so this is the normal growth rate.”
Plan. Check the slope’s units and causal wording, translate \(r^2\) into variation accounted for, and check whether zero fertilizer is within the observed range before accepting the practical interpretation of the intercept.
Do. The slope’s units are centimeters of growth per month for each additional tablespoon of fertilizer per week. Since amounts were chosen by the students rather than assigned in an experiment, the data show an association, not that fertilizer caused the predicted increase. The phrase “fertilizer explains 64% of growth” is too vague: \(0.64\times100\%=64\%\) of the variation in monthly plant growth among these plants is accounted for by the linear regression using fertilizer amount.
The line predicts \(1.2\) centimeters of growth per month at zero tablespoons per week. But zero is below the observed range of 0.5 to 3.0 tablespoons per week, so that interpretation extends the line beyond the observed fertilizer amounts. Calling 1.2 centimeters the “normal growth rate” also suggests a general or typical rate that these data do not establish.
Conclude. A careful rewrite is: “Among these 30 basil plants in the greenhouse, each additional tablespoon of fertilizer per week was associated with a predicted increase of 2.5 centimeters in monthly growth. About 64% of the variation in monthly growth among these plants is accounted for by the fitted linear model. Because fertilizer amounts were chosen by the students and the observed amounts were 0.5 to 3.0 tablespoons per week, these results do not establish a causal effect, and the intercept at zero tablespoons has limited practical meaning.”
Do Not Let a Large \(r^2\) Become a Stronger Claim
A high \(r^2\) may make an argument sound persuasive, but it does not repair missing context or justify a causal conclusion. It describes the proportion of variation in the response accounted for by the linear model in the data used to fit it. It does not tell us that the explanatory variable caused that variation, that every prediction is close, or that the model applies to different cases or settings.
Consider a second kind of overstatement: treating a precise-looking fitted value as an observed fact. In an invented delivery-route analysis, 42 routes had distances from 2 to 18 kilometers and delivery times in minutes. The fitted line was \(\hat{y}=11+3.8x\), with \(r^2=0.92\) and \(s=6.5\) minutes. A writer says, “Each extra kilometer makes every delivery take 3.8 minutes longer, and distance determines 92% of delivery time. A 20-kilometer delivery takes 87 minutes.”
The line’s value at 20 kilometers is \(11+3.8(20)=87\) minutes, but 20 kilometers is beyond the observed range of 2 to 18 kilometers, so this is an extrapolation. The slope describes a predicted change, not the change for every individual route or proof that distance caused it. The \(r^2\) interpretation should refer to variation in delivery times accounted for by the linear regression using distance, not to a percentage “determined” by distance. The reported \(s\) is 6.5 minutes, which helps describe typical residual size; it does not guarantee that any particular prediction will be within 6.5 minutes.
The collection method matters, too. If routes with longer distances also tend to have different traffic conditions or delivery procedures, those features may be relevant to the observed association. Unless the data-collection design supports a causal conclusion, a reviewer should recommend wording about association rather than cause and effect. Regression output alone does not settle that question.
Worked Example: Critique a Complete but Misleading Summary
Situation. An invented analysis uses data from 42 delivery routes operated by one local service during a two-week period. Distance in kilometers is \(x\), and delivery time in minutes is \(y\). The observed distances range from 2 to 18 kilometers. The fitted line is \(\hat{y}=11+3.8x\), with \(r^2=0.92\) and \(s=6.5\) minutes.
A report says: “The model proves distance is the reason deliveries take longer. Ninety-two percent of delivery times are explained correctly, and the company can use the equation for all its routes.”
State. The report misstates \(r^2\), makes a causal claim unsupported by the described data, and applies the model beyond the setting represented by the cases without providing evidence for that wider scope.
Plan. Replace “proves” and “the reason” with a description of association. Correct the meaning of \(r^2\), and limit the conclusion to the observed routes and setting unless additional evidence supports applying the model more widely.
Do. Since \(0.92\times100\%=92\%\), a correct interpretation is: “About 92% of the variation in delivery times among the 42 routes is accounted for by the linear regression of delivery time on route distance.” This does not mean 92% of times are predicted correctly. The \(s\) value indicates that residuals typically differ from the fitted line by about 6.5 minutes in this context.
Conclude. A stronger report would say: “For the 42 routes operated by this service during the two-week period, delivery distance and delivery time had a positive linear association. The fitted model accounts for about 92% of the variation in delivery times among these routes, with a typical residual size of 6.5 minutes. These data do not by themselves show that distance causes longer delivery times or establish that the model applies to every route operated by other services or in other settings.”
Common Mistakes and AP Exam Tips
- Criticizing without naming the problem. “The conclusion is wrong” does not show what is wrong. Identify the phrase, such as “predicted correctly,” explain the correct meaning of the statistic, and connect that correction to the situation.
- Fixing a number but leaving out the context. Changing “64%” to “64% of variation” is not yet a complete interpretation. Say variation in which response, among which cases, accounted for by which regression model.
- Confusing a predicted change with a causal effect. A slope describes predicted change in the fitted relationship. Unless the data-collection design supports causal inference, describe the association and avoid saying that changing \(x\) will cause \(y\) to change.
- Ignoring scope when a sentence sounds broad. A result from one school, greenhouse, company, or time period does not automatically apply to all students, plants, companies, or future periods. Name the cases and setting represented in the data.
- Not checking the explanatory-variable range. A predicted response can be mathematically calculated outside the observed \(x\)-range, but that does not make it dependable. Follow the range-checking approach from “Writing an Extrapolation Critique”: identify the requested \(x\)-value, compare it with the observed range, and explain the implication.
- Rewriting more strongly than the evidence allows. A critique should not replace an overclaim with another overclaim. Use “is associated with” for an observed relationship and reserve causal language for situations where the design supports it.
For full-credit communication, make the criticism specific and provide a defensible alternative. A clear answer might say: “The statement that 64% of plants were predicted correctly is an incorrect interpretation of \(r^2\). The value means that 64% of the variation in monthly plant growth among the observed plants is accounted for by the linear regression using fertilizer amount.” When the issue is causation, explain that the data show an association and state why the described collection method does not establish cause and effect.
Check Your Understanding
For each item, identify the problem and write a more defensible version of the statement.
- A regression of weekly screen time on sleep duration has \(r^2=0.49\). A student writes, “The model predicts 49% of teenagers correctly.” What does \(r^2\) describe instead, and what context should the interpretation include?
- An observational project finds a positive slope between hours of practice and performance score. The report says, “Practice causes higher scores.” What is the overclaim, and what wording is supported by the described design?
- A line predicts household electricity use from the number of appliances. The report interprets its intercept as typical electricity use, but the observed appliance counts range from 3 to 12. What should the reviewer check and explain?
- A model is fitted to deliveries from one town during one month. The conclusion says it applies to all delivery services. Identify the missing scope limitation.
- A writer reports a slope as “2.1” without naming either variable or its units. What information should a complete slope interpretation provide?