Comparing Two Linear Models for the Same Response
In “What \(r^2\) Does Not Tell You,” you learned that \(r^2\) summarizes the proportion of variation in a response accounted for by its linear relationship with an explanatory variable. This tutorial uses that summary for a specific purpose: choosing between two linear models for the same response.
Suppose a community group wants to predict weekly water use for households. It fits one linear model using household size and another using lot area as the explanatory variable. If both models use the same water-use measurements for the same households, their \(r^2\) values can be compared. The model with the larger \(r^2\) accounts for a greater proportion of the variation in water use through its linear relationship with the explanatory variable.
That is a focused comparison, not a complete verdict about which model is best for every purpose. As the earlier tutorial explains, \(r^2\) alone does not show whether the pattern is suitably linear, whether the model’s predictions are close for individual cases, or whether an explanatory variable causes the response to change.
Make Sure the Comparison Is Fair
A meaningful comparison begins by checking that the models are being judged on the same response values. The observations should be the same, too. If one model is fitted to 40 households and another to a different 32 households, their \(r^2\) values summarize variation in different data sets. A difference could reflect which cases were included rather than a genuine difference between the explanatory variables.
The response must also be measured in the same way for both models. For example, comparing two models for the same students’ exam scores is different from comparing one model for exam scores with another for study hours. Even if both \(r^2\) values are reported, they describe variation in different responses and do not answer which model accounts for more variation in the same response.
Once the response and observations match, compare the \(r^2\) values directly. You can express each as a percentage by multiplying by 100. If you subtract the smaller percentage from the larger one, describe the result as a difference in percentage points. For instance, a comparison of 60% and 45% is a difference of 15 percentage points—not a difference of 15%.
Confirm that both models predict the same response variable, recorded in the same units and for the same cases.
Identify which value is larger. Convert to percentages if that makes the comparison clearer.
Explain which model accounts for more of the response’s variation through its linear relationship, and quantify the difference when useful.
Say that one model is better by the \(r^2\) criterion. Do not treat the comparison as proof of causation, appropriate linearity, or accurate predictions for every case.
Worked Comparisons
Worked Example: Choosing Between Two Models of Water Use
A fictional community group records weekly water use, in hundreds of liters, for the same 30 households. It fits one linear model using household size and another using lot area. The first model has \(r^2=0.64\), and the second has \(r^2=0.49\). Which model accounts for more variation in weekly water use?
State. The response is weekly household water use, measured in hundreds of liters. The explanatory variables are household size and lot area.
Plan. First confirm that the models use the same response values and the same households. Since both models use weekly water use for all 30 households, compare their \(r^2\) values directly. The larger \(r^2\) identifies the model that accounts for more response variation by this measure.
Do. The household-size model has \(r^2=0.64\), or \(0.64(100\%)=64\%\). The lot-area model has \(r^2=0.49\), or \(0.49(100\%)=49\%\). The difference is \(64\%-49\%=15\) percentage points. Since \(0.64>0.49\), the household-size model has the larger \(r^2\).
Conclude. For these 30 households, about 64% of the variation in weekly water use is accounted for by its linear relationship with household size, compared with about 49% accounted for by its linear relationship with lot area. The household-size model accounts for 15 percentage points more of the variation, so it is the better choice of these two by the \(r^2\) criterion. This comparison alone does not show that household size causes water use to change or that the model’s predictions are close for every household.
Worked Example: Reading a Small Difference Correctly
A fictional school analyzes the same 48 students’ science scores using two explanatory variables: time spent on a practice website and number of science classes taken. Both models use science score, measured in points, as the response. The practice-time model has \(r^2=0.72\); the classes model has \(r^2=0.68\). A student says the first model is “72% accurate” and is “4% better.” Evaluate that statement and make the comparison.
State. Both models have the same response—science score in points—and use the same 48 students. Therefore, their \(r^2\) values can be compared.
Plan. Interpret each \(r^2\) as a proportion of variation in science scores accounted for by the model’s linear relationship with its explanatory variable. Convert the difference to percentage points. Do not describe \(r^2\) as the percentage of accurate predictions.
Do. For practice time, \(0.72(100\%)=72\%\) of the variation in science scores is accounted for by the linear relationship with practice time in these data. For number of science classes, \(0.68(100\%)=68\%\) is accounted for by its linear relationship with science scores. The difference is \(72\%-68\%=4\) percentage points.
Conclude. The practice-time model accounts for 4 percentage points more of the variation in science scores than the classes model, so it is better by the \(r^2\) criterion. The statement “72% accurate” is incorrect: \(r^2=0.72\) does not mean that 72% of individual predictions are accurate. Also, the difference is 4 percentage points, not automatically “4% better.” The two models’ \(r^2\) values describe how much variation each accounts for, not the proportion of predictions that meet some accuracy standard.
Worked Example: Why Different Sets of Cases Complicate the Choice
A fictional environmental team models the same response, daily stream depth in centimeters. A model using rainfall has \(r^2=0.81\) for 25 days with complete rainfall records. A model using temperature has \(r^2=0.86\) for 19 different days with complete temperature records. Can the team conclude that the temperature model is better because \(0.86>0.81\)?
State. The response is daily stream depth in centimeters. Although the response variable is the same, the models were fitted to different sets of days.
Plan. Check comparability before applying the larger-\(r^2\) rule. Since the models do not use the same response observations, the values summarize variation in different data sets. A direct comparison cannot isolate which explanatory variable is associated with more variation for the same cases.
Do. The rainfall model’s \(r^2=0.81\) means that about 81% of the variation in stream depth among its 25 included days is accounted for by its linear relationship with rainfall. The temperature model’s \(r^2=0.86\) means that about 86% of the variation among its 19 included days is accounted for by its linear relationship with temperature. The numerical difference is \(0.86-0.81=0.05\), or 5 percentage points, but those percentages come from different sets of days.
Conclude. The team cannot fairly conclude from these values alone that the temperature model is better. Its larger \(r^2\) describes the 19 days used for that model, while the other value describes 25 different days. For a direct \(r^2\) comparison, the team should fit both models using the same days and the same stream-depth measurements, then compare the resulting values.
What a Larger \(r^2\) Does—and Does Not—Support
When the comparison is fair, a larger \(r^2\) supports a precise descriptive conclusion: in the data being compared, that model accounts for a larger proportion of the response’s variation through its linear relationship with the explanatory variable. It is reasonable to call it the better of the two by \(r^2\). That qualifier matters because model choice can involve more than this one summary.
A larger \(r^2\) does not automatically establish that the linear model is appropriate. In “Curved Patterns in a Residual Plot,” you learned that a systematic curve in the residuals indicates that a line misses in an organized way. If one model has the larger \(r^2\) but its residual plot shows a clear curved pattern, the larger value does not erase that concern. Check the scatterplot and residual plot when evaluating whether a linear model is suitable.
Nor does the larger \(r^2\) necessarily mean that the model has smaller prediction errors for every case or is more useful for a particular decision. As discussed in “Using \(s\) to Describe Prediction Accuracy,” the residual standard deviation \(s\) gives information about the typical size of prediction errors in the response’s units. A comparison of \(r^2\) answers how much response variation is accounted for; it does not replace consideration of residuals, error scale, or the purpose of the prediction.
Finally, \(r^2\) does not establish causation. If the data come from an observational study, selecting the model with the larger \(r^2\) does not show that its explanatory variable causes the response to change. Describe association, and let the study design determine whether causal conclusions are justified.
Common Mistakes and AP Exam Tips
- Comparing models that use different observations. Before choosing the larger value, state whether both models use the same response values for the same cases. If they do not, explain why the \(r^2\) comparison is not direct.
- Comparing models with different responses. A larger \(r^2\) for one response does not establish a better model for another response. Name the response in your interpretation.
- Calling the difference a percent when it is percentage points. If one model accounts for 72% and the other for 68%, the difference is 4 percentage points. Show the subtraction when reporting the difference.
- Calling \(r^2\) “percent accurate.” Full-credit wording describes the percentage of variation in the response accounted for by the linear relationship. It does not describe the percentage of predictions that are correct.
- Claiming the larger \(r^2\) proves the model is best overall. Say “better by the \(r^2\) criterion” and, when appropriate, mention that residual patterns and prediction error also matter.
- Turning association into causation. A larger \(r^2\) does not show that the explanatory variable causes the response to change. Keep the conclusion descriptive unless the study design supports a causal claim.
A clear AP response identifies the common response, confirms that the comparison uses the same observations, interprets both values in context, and states which model has the larger \(r^2\). If you report the gap, use percentage points. Then limit the conclusion: the selected model accounts for more variation by this measure, not necessarily that it is appropriate or superior in every other respect.
Check Your Understanding
Use the comparison rule and its limitations to answer each question.
- Two models use the same 36 garden plots and predict tomato yield in kilograms. Their \(r^2\) values are 0.55 and 0.70. Which model accounts for more variation, and by how many percentage points?
- A model of commute time has \(r^2=0.64\). Explain what this value means without calling it the percent of accurate predictions.
- Two models predict the same response, but one uses 50 observations and the other uses 42 different observations. Why is choosing the larger \(r^2\) not a fair direct comparison?
- Two comparable models have \(r^2\) values of 0.83 and 0.76. What does the larger value support, and what does it not establish about causation?
- The model with the larger \(r^2\) has a curved residual pattern. Why should that pattern still matter when deciding whether to use a linear model?