Two Possible Predictors, One Response
Suppose a teacher wants to predict students’ exam scores and has recorded both hours studied and hours slept the night before the exam. Each variable could be used as the explanatory variable in a separate simple linear regression model. The question is not which variable has the steeper fitted line; it is which model describes the observed scores more effectively for the intended purpose.
As in “Comparing Two Linear Models for the Same Data,” compare the coefficient of determination, \(r^2\), and the residual standard deviation, \(s\). For this comparison to be fair, both regressions should use the same response values—the same students’ exam scores—and the same students. If that is true, the model with larger \(r^2\) also has smaller \(s\). The two summaries describe fit in complementary ways: \(r^2\) is a proportion of response variation, while \(s\) is measured in the response’s units.
What to Read in the Regression Output
A regression output may report the fitted line, \(r^2\), \(s\), and the number of observations. The fitted line shows the predicted score for each value of its explanatory variable. Its slope is measured in score points per hour, but a slope describes the rate of change in predicted score for a one-hour increase in that particular predictor. Because hours studied and hours slept can have different relationships with score and different spreads among students, comparing the slopes does not rank the models’ fit.
Instead, compare \(r^2\) and \(s\). A model with a larger \(r^2\) accounts for a larger proportion of the variation in the observed scores. A model with a smaller \(s\) has smaller typical residuals, measured in score points. The earlier tutorial “Choosing Numbers to Support a Linear Model” explains these summaries; here, the goal is to use them side by side for two different predictors.
The numerical comparison is not the whole decision. “Choosing Graphs to Support a Linear Model” and “Reading a Residual Plot for Model Fit” explain why a scatterplot and residual plot matter: a higher \(r^2\) does not make a curved relationship linear or rule out changing spread. Also consider whether the predictor is available for the prediction being planned. The comparison describes association, not cause and effect.
- Both models use the same response variable, measured on the same scale.
- Both models use the same cases, so each score is included in both fits.
- Both are simple linear regression models with an intercept.
- Use the scatterplots and residual plots to assess whether a linear model is a reasonable description for each predictor.
These are checks for a useful, fair comparison, not a set of conditions for a significance test. In this tutorial, the goal is to describe which model fits the observed scores better, not to test whether one predictor has a statistically significant effect.
A Clear Comparison Process
Use the following sequence when an output gives separate regressions for hours studied and hours slept. It keeps the comparison focused on the task and helps prevent a tempting but unsupported conclusion.
Confirm that both outputs use exam score as the response and the same students. Check the reported sample sizes and, if available, the case identifiers.
Identify the larger \(r^2\) and smaller \(s\). State what each means for scores, using score points for \(s\).
Look for patterns in each scatterplot and residual plot. Consider whether the predictor will be available and appropriate when scores need to be predicted.
Say which model is favored by the fit summaries for the students observed, and mention any important limitation. Do not claim that the predictor causes score differences.
Worked Example: Compare the Two Outputs
Worked Example: Hours Studied or Hours Slept?
Original AP-style question. In an invented classroom data set, a teacher fits two regressions using the same 20 students. In both models, exam score \(y\) is measured in points, from 0 to 100. The observed scores have \(SST=2400\text{ points}^2\). Use the output to decide which variable predicts score better in these data.
| Explanatory variable | Fitted line | \(r^2\) | \(s\) | \(n\) |
|---|---|---|---|---|
| Hours studied | \(\hat{y}=59.4+5.20x\) | 0.640 | 6.93 points | 20 |
| Hours slept | \(\hat{y}=20.4+7.80x\) | 0.360 | 9.24 points | 20 |
State. We will compare the hours-studied and hours-slept models for predicting the same 20 students’ exam scores.
Plan. The response and the students are the same in both fits, the response is measured in points in both, and both models are simple linear regressions with intercepts. Those facts make the comparison fair. The table gives \(r^2\) and \(s\), so we can compare the proportion of score variation accounted for and the typical size of the residuals. We would also inspect the scatterplots and residual plots before relying on either line.
Do. The hours-studied model has \(r^2=0.640\), so it accounts for 64.0% of the variation in the observed exam scores. The hours-slept model has \(r^2=0.360\), so it accounts for 36.0%. The residual standard deviations are 6.93 points and 9.24 points, respectively. As a check using the supplied \(SST\), the first model has \(SSE=(1-0.640)(2400)=864\text{ points}^2\), so \(s=\sqrt{864/(20-2)}\approx6.93\) points. For hours slept, \(SSE=(1-0.360)(2400)=1536\text{ points}^2\), so \(s=\sqrt{1536/18}\approx9.24\) points. The calculations agree with the reported output.
Conclude. For these 20 students, the hours-studied model is favored by both fit summaries: it accounts for a greater proportion of score variation, and its residuals are typically smaller by this response-scale measure. This does not prove that studying caused higher scores. Before using the model for predictions, check the plots and consider whether the observed students and predictor values are relevant to the intended use.
Notice that the hours-slept slope, 7.80 score points per hour, is steeper than the hours-studied slope, 5.20 points per hour. That does not make the sleep model better: its \(r^2\) is smaller and its \(s\) is larger. Slope describes the fitted rate of change, not the overall closeness of the data to the line.
Worked Example: Different Sample Sizes Need a Fairer Check
Worked Example: Missing Sleep Records
Original AP-style question. In a different invented class, the hours-studied output uses 20 students, while the hours-slept output uses only 17 because three students did not report sleep hours. The outputs initially suggest \(r^2=0.58\) for study and \(r^2=0.62\) for sleep. What can be concluded, and what should the analyst do?
Solution. The initial outputs do not use the same cases. The \(r^2\) values describe fits to different sets of students, so their numerical difference does not provide a clean comparison of the two predictors on the same scores. The analyst should fit both models using the same 17 students with recorded sleep hours, rather than comparing the original values as if the samples matched.
Suppose the refitted outputs on those same 17 students report \(SST=900\text{ points}^2\), \(r^2=0.4900\) for hours studied, and \(r^2=0.5625\) for hours slept. Both fits have \(n=17\), so each has \(17-2=15\) residual degrees of freedom. For study, \(SSE=(1-0.4900)(900)=459\text{ points}^2\), giving \(s=\sqrt{459/15}=\sqrt{30.6}\approx5.53\) points. For sleep, \(SSE=(1-0.5625)(900)=393.75\text{ points}^2\), giving \(s=\sqrt{393.75/15}=\sqrt{26.25}\approx5.12\) points. These calculations agree with the shared \(SST\) and sample size.
Conclusion. On the common set of 17 students, the hours-slept model has the larger \(r^2\) and smaller \(s\), so it fits those observed scores better according to these summaries. The conclusion applies to this common subset. Because students without sleep records were excluded, consider whether that omission could limit how well the comparison represents the full class.
Worked Example: A Small Difference and a Steeper Slope
Worked Example: Avoid Choosing by Slope Alone
Original AP-style question. For the same 15 students, two invented regression outputs predict exam score in points. The study model has slope 4.10 points per additional study hour and \(r^2=0.5184\). The sleep model has slope 6.80 points per additional sleep hour and \(r^2=0.5041\). Both use scores with \(SST=1400\text{ points}^2\). Compare the models, including their residual standard deviations.
Solution. The models use the same response and students, so their fit summaries can be compared. The study model accounts for 51.84% of the score variation, while the sleep model accounts for 50.41%. The residual degrees of freedom are \(15-2=13\). For study, \(SSE=(1-0.5184)(1400)=674.24\text{ points}^2\), and \(s=\sqrt{674.24/13}\approx7.20\) points. For sleep, \(SSE=(1-0.5041)(1400)=694.26\text{ points}^2\), and \(s=\sqrt{694.26/13}\approx7.31\) points. The study model has slightly smaller typical residuals, consistent with its slightly larger \(r^2\).
Conclusion. The output gives hours studied a small numerical advantage, not a decisive one. Although the sleep slope is steeper, that reflects a larger fitted change in score per hour on its own scale; it does not show that sleep is the better predictor. If sleep hours are more reliably recorded or are available earlier for a particular planning task, that practical consideration may matter, but it does not change which model has the slightly better fit summaries in these data. Examine both residual plots and avoid presenting a small difference as proof that one predictor is universally better.
Common Mistakes and AP Exam Tips
- Choosing the steeper line. The predictors have different scales and may have different spreads. Compare \(r^2\) and \(s\) to discuss fit; interpret the slope only as the predicted score change per additional hour of that specific predictor.
- Comparing regressions with different cases. Different sample sizes may mean different students were used. State the problem and, when possible, compare refitted models using the same cases.
- Calling \(r^2\) the percent of scores predicted correctly. Say that the model accounts for a stated percentage of the variation in observed scores. It is not a percentage of individual predictions that are correct.
- Leaving off the units for \(s\). Because the response is exam score, \(s\) is measured in score points. Describe it as a typical residual size, not a guaranteed maximum prediction error.
- Claiming causation. A regression comparison shows how the predictors are associated with scores in the observed data. Unless the study design supports a causal conclusion, do not say that more sleep or study time caused a score change.
- Treating a small numerical edge as decisive. Report the size of the difference and check the plots, data scope, and prediction goal. A slightly higher \(r^2\) does not settle every practical choice.
A full-credit comparison names the response and confirms that the same cases were used, compares \(r^2\) and \(s\) with appropriate interpretations and units, and makes a qualified conclusion in context. If the samples differ, explain why the direct comparison is not fair and identify the need for a common set of cases.
Check Your Understanding
Use the regression-comparison ideas from this tutorial to answer each question.
- Two models predict the same exam scores, but one uses 20 students and the other uses 18. Why is comparing their \(r^2\) values directly potentially unfair?
- For the same 24 students, the study model has \(r^2=0.70\) and \(s=5.2\) points. The sleep model has \(r^2=0.45\) and \(s=7.8\) points. Which model is favored by these summaries, and what does its \(r^2\) mean?
- The sleep model has a slope of 8 points per hour, while the study model has a slope of 5 points per hour. Does that establish that sleep predicts score better? Explain.
- Why should each residual standard deviation be reported in score points?
- Name one graphical check and one practical consideration to consider before using the favored model for prediction.