Start With the Claim, Then Choose the Evidence
A regression report may contain an equation, \(r\), \(r^2\), and \(s\), while the analysis may also include a scatterplot and a residual plot. A strong written conclusion does not simply repeat all of these. It selects the evidence that directly supports the claim being made.
In “Choosing Numbers to Support a Linear Model” and “Choosing Graphs to Support a Linear Model,” you learned what these values and graphs describe. Here the task is to match them to a particular sentence. For example, \(r\) describes the direction and strength of a linear association. A residual plot helps assess whether the fitted line leaves a systematic pattern in its errors. Those pieces of evidence are related, but they do not mean the same thing.
Before writing, identify the central verb or idea in the claim. Is the claim about an association being positive or negative, about its strength, about whether a straight-line model is reasonable, about the proportion of response variation accounted for, or about the typical size of prediction errors? Each question calls for different evidence.
| Claim being considered | Evidence that directly addresses it | What the evidence does not establish by itself |
|---|---|---|
| The linear association is positive or negative. | The sign of \(r\), together with the direction in the scatterplot. | That \(x\) causes changes in \(y\). |
| The linear association is strong or weak. | The magnitude of \(r\), considered with the scatterplot. | That a straight-line model captures every feature of the data. |
| A linear model is reasonable for the observed pattern. | The scatterplot and a residual plot without a clear systematic pattern. | That the model will work equally well outside the observed range. |
| The model accounts for a stated proportion of response variation. | \(r^2\), interpreted as a proportion of variation in the observed responses. | That the same percentage of cases is predicted exactly. |
| Predictions typically differ from observed responses by a certain amount. | The residual standard deviation, \(s\), in response units. | That every individual prediction misses by exactly that amount. |
| The fitted response changes at a stated rate as \(x\) increases. | The slope, interpreted in context with its units. | That the relationship is causal or that every case follows the slope exactly. |
A good conclusion usually needs only the evidence relevant to the question. If asked whether there is a strong positive linear association, citing the equation’s intercept is beside the point. If asked whether the fitted line is appropriate, citing a high \(r\) alone is incomplete: a residual plot might still reveal curvature or changing spread, as discussed in “Reading a Residual Plot for Model Fit.”
Build a Claim-and-Evidence Sentence
A useful approach is to make the claim, name the evidence, and state how the evidence supports the claim. Include the variables or their context so the reader knows what the statistics describe. If a graph is part of the evidence, describe the visible feature rather than saying only that the graph “looks good.”
Underline the specific idea: direction, strength, linear fit, explained variation, typical error, or rate of change.
Choose the statistic or graph that measures or displays that idea. Do not substitute a convenient but unrelated number.
State what the value or graph shows, using the variables, context, and units where appropriate.
Do not claim more than the evidence supports. For example, a strong association is not proof of causation, and a patternless residual plot is not a guarantee of accurate predictions.
These steps help avoid two opposite problems. One is an unsupported conclusion, such as calling a model “a good fit” without describing any evidence about the residuals. The other is a list of statistics without an argument: reporting \(r\), \(r^2\), and \(s\) is not enough unless the answer explains which value supports which claim.
Worked Example: Evidence for a Strong Positive Linear Association
Worked Example: Trail Use and Daily Temperature
Original AP-style question. In an invented set of observations from 12 days, a park planner records the afternoon temperature in degrees Celsius and the number of trail entries that day. The scatterplot shows an upward, approximately straight pattern. The regression report gives \(r=0.93\), \(r^2=0.8649\), and \(s=18.6\) entries. The residual plot shows points scattered around zero with no clear curve or change in spread. A student writes, “There is a strong positive linear association between afternoon temperature and trail entries, and a linear model is reasonable for these observed days.” Identify the evidence that supports each part of the statement.
State. Evaluate separately the claims about association and about the suitability of a linear model.
Plan. The direction and strength of the linear association are addressed by the scatterplot and the sign and magnitude of \(r\). The residual plot is relevant to whether a straight-line model leaves a systematic pattern. The values \(r^2\) and \(s\) describe other features, so they are not the primary evidence for either of these two claims.
Do. The positive value \(r=0.93\) is close to 1, so it supports describing the linear association between afternoon temperature and trail entries as strong and positive. The upward, approximately straight pattern in the scatterplot is consistent with that description. The value \(r^2=0.8649\) is consistent with the reported correlation because \(0.93^2=0.8649\), but \(r^2\) is not needed to establish the direction of the association.
The residual plot has points scattered around zero without a clear curved pattern or changing spread. That visible feature supports using a linear model to summarize the pattern in these observations. It is evidence about the model’s residual behavior, not proof that the observations were collected randomly; “random-looking residuals” describes the absence of an obvious pattern in the plot.
Conclude. For the 12 observed days, the upward scatterplot and \(r=0.93\) support a strong positive linear association between afternoon temperature and trail entries. The residuals show no clear systematic pattern, which supports using a linear model for these observations. This evidence does not show that higher temperature causes more trail use.
Worked Example: A High Correlation Does Not Settle Model Fit
Worked Example: Rainfall and Creek Turbidity
Original AP-style question. In an invented monitoring activity, rainfall amount \(x\), in centimeters, and creek turbidity \(y\), in a measurement unit used by the activity, are recorded for 15 rain events. The scatterplot shows a clear increasing trend. A regression report gives \(r=0.95\), \(r^2=0.9025\), and \(s=4.8\) turbidity units. However, the residual plot has a fan shape: residuals are more tightly grouped at smaller fitted turbidity values and more spread out at larger fitted values. A draft conclusion says, “Because \(r=0.95\), the linear model fits the data well and predicts turbidity equally reliably throughout the range.” Revise the conclusion using evidence matched to each claim.
Solution. The positive correlation \(r=0.95\), together with the increasing scatterplot, supports a strong positive linear association between rainfall and turbidity for these observed events. The coefficient of determination is consistent with the correlation because \(0.95^2=0.9025\). It means 90.25% of the variation in observed turbidity values is accounted for by the fitted linear model; it does not show that predictions are equally reliable at every rainfall level.
The fan-shaped residual plot is the important evidence for the claim about uniform fit. Because the residual spread increases at larger fitted turbidity values, the model’s errors vary in size across the range. Thus, the large \(r\) supports a strong linear association, but it does not justify saying that the model fits equally well or predicts equally reliably throughout the range. The reported \(s=4.8\) turbidity units summarizes typical residual size overall; it does not erase the changing spread visible in the residual plot.
A more defensible conclusion is: “Rainfall and turbidity have a strong positive linear association in these 15 observed events (\(r=0.95\)). However, the residual plot shows increasing spread at larger fitted values, so the size of prediction errors is not uniform across the range.” This conclusion retains the claim supported by \(r\) and qualifies the fit claim using the residual evidence.
Worked Example: Select Evidence for Several Different Claims
Worked Example: Delivery Distance and Travel Time
Original AP-style question. An invented delivery service records route distance \(x\), in kilometers, and travel time \(y\), in minutes, for 20 deliveries. The fitted line is \(\hat{y}=11.2+3.4x\). The report gives \(r=0.88\), \(r^2=0.7744\), and \(s=6.2\) minutes. The scatterplot has an increasing, approximately straight pattern, and the residual plot shows a roughly even band of points around zero without an obvious pattern. For each claim below, choose the most relevant evidence and write an appropriately qualified sentence.
- For each additional kilometer, the model predicts a certain increase in travel time.
- The linear association is strong and positive.
- The model accounts for a substantial proportion of the variation in travel time.
- Predicted travel times have a typical residual size measured in minutes.
- A linear model is reasonable for describing the observed pattern.
Solution. The slope \(3.4\) directly supports the first claim. In context, for each additional kilometer of route distance, the fitted model predicts an increase of 3.4 minutes in travel time, on average across the line. The slope’s units are minutes per kilometer. This describes the fitted relationship; it does not mean every additional kilometer adds exactly 3.4 minutes to every delivery.
The correlation \(r=0.88\) and the upward scatterplot support the second claim. Since \(r\) is positive and its magnitude is fairly close to 1, they indicate a strong positive linear association between route distance and travel time in these 20 deliveries. The residual plot is not the main evidence for direction or strength, though it adds information about how the line’s errors behave.
For the third claim, use \(r^2=0.7744\). In context, 77.44% of the variation in the observed travel times is accounted for by the fitted linear model. This does not mean that 77.44% of deliveries have exactly correct predictions.
For the fourth claim, use \(s=6.2\) minutes. It indicates that observed travel times typically differ from their fitted values by about 6.2 minutes. The unit is minutes because \(s\) is measured in the response’s units. It is a typical scale, not the exact error for every delivery.
For the final claim, use the scatterplot and residual plot together. The scatterplot is approximately straight, and the residual plot shows a roughly even band around zero without an obvious pattern. These features support a linear model for describing the observed deliveries. They do not guarantee that the model is appropriate outside the observed route-distance range or for deliveries in a different setting.
Common Mistakes and AP Exam Tips
- Using \(r\) as proof that a line fits well. A high \(|r|\) supports a strong linear association, but inspect the residual plot for systematic structure or changing spread before making a claim about fit.
- Using \(r^2\) to describe strength or direction. \(r^2\) is a proportion of variation accounted for by the model and has no sign. Use \(r\) and the scatterplot when the claim is about direction and strength.
- Calling a residual plot “random” without saying what it shows. A stronger answer says the residuals are scattered around zero with no clear curve or other systematic pattern. This is a description of the plot, not a claim that data collection was random.
- Using \(s\) to claim that predictions are equally accurate everywhere. \(s\) gives a typical residual size in response units. Check the residual plot for whether the spread changes across the fitted values.
- Listing every statistic without linking it to the conclusion. Choose the values that answer the question and explain their meaning in context. Extra numbers do not replace a reasoned connection between evidence and claim.
- Overstating what a model summary proves. Do not turn association into causation, claim that predictions are exact, or generalize beyond the cases and range supported by the data.
For full credit, name the evidence and interpret it rather than merely quoting a number. For instance, “\(r=0.88\)” alone is incomplete. A stronger sentence is: “For these deliveries, \(r=0.88\) and the scatterplot slopes upward, supporting a strong positive linear association between route distance and travel time.” For model fit, describe the residual pattern separately. Keeping these claims distinct makes the conclusion both clearer and more accurate.
Check Your Understanding
For each situation, identify evidence that directly supports the claim and state one limitation of that evidence.
- A conclusion says two variables have a strong negative linear association. Which regression value and graph feature would you use to support that statement?
- A report gives \(r=0.97\), but the residual plot shows a clear curve. What does the correlation support, and what does the residual pattern suggest about using a straight line?
- A student says that \(r^2=0.81\) means 81% of cases are predicted exactly. Correct the claim using the appropriate interpretation of \(r^2\).
- A model report gives \(s=5.4\) points, and the response is a test score. What claim does \(s\) support, and what should not be inferred from it?
- A scatterplot is approximately straight and a residual plot shows a roughly even band around zero. Write one sentence describing how this evidence supports a linear model, with an appropriate qualification.