Tutorials › AP Statistics › Comparing Two Linear Models for the Same Data

Comparing and communicating regression models · Tutorial 984 of 1000

Comparing Two Linear Models for the Same Data

Compare how much response variation two candidate predictors account for and how large their typical residuals are, then justify which model is more useful in context.

Intermediate 9 min read

What You'll Learn

  • Compare \(r^2\) and \(s\) when two simple linear models use the same response values and cases.
  • Explain why, under that fair comparison, a larger \(r^2\) corresponds to a smaller \(s\).
  • Calculate and interpret \(r^2\) and \(s\) from residual information.
  • Use residual plots and context alongside numerical model-fit summaries.
  • Justify a model choice without claiming that association proves causation.

Comparing Models on a Fair Basis

A data set can include several quantitative variables that might help predict the same response. For example, a gardener might consider both fertilizer amount and watering frequency as predictors of plant growth. If each candidate predictor is used in a separate least-squares regression line, how should we compare the models?

This tutorial focuses on two numerical summaries from “Choosing Numbers to Support a Linear Model”: the coefficient of determination, \(r^2\), and the residual standard deviation, \(s\). The first describes the proportion of response variation accounted for by the linear model; the second describes the typical size of residuals in response units. To compare them fairly, fit both models to the same response values for the same cases.

Definition: A candidate-predictor comparison fits separate simple linear regression models to the same observed response values, using a different quantitative explanatory variable in each model. Comparing \(r^2\) and \(s\) helps describe which fitted line accounts for more response variation and has smaller typical residuals.

A comparison is not fair if one model uses a different response variable, a different set of observations, or a different response scale. Those changes can affect the summaries independently of how well a predictor fits. State which response and cases are being compared, and identify the explanatory variable in each model.

Why \(r^2\) and \(s\) Agree in This Comparison

For a simple linear regression with an intercept, let \(SST=\sum (y-\bar{y})^2\) be the total sum of squares in the response and let \(SSE=\sum (y-\hat{y})^2\) be the sum of squared residuals. The coefficient of determination is \(r^2=1-SSE/SST\). The residual standard deviation is \(s=\sqrt{SSE/(n-2)}\), where \(n\) is the number of observations.

When both models use the same response values and cases, \(SST\) and \(n\) are unchanged. So a larger \(r^2\) means a smaller \(SSE\), and a smaller \(SSE\) means a smaller \(s\). The two summaries express the same ordering of fit in different ways: \(r^2\) as a proportion of response variation, and \(s\) in response units.

$$ r^2=1-\frac{SSE}{SST}, \qquad s=\sqrt{\frac{SSE}{n-2}} $$
Key comparison: With the same response values, the same cases, and a simple linear model with an intercept for each predictor, the model with the larger \(r^2\) also has the smaller \(s\). This does not by itself establish that the model is useful, that its pattern is linear, or that its predictor causes changes in the response.

Even when the rankings agree, report what each number means. A higher \(r^2\) says a larger proportion of the variation in the observed response values is accounted for by that linear model. A lower \(s\) says the observed responses are typically closer to that model’s fitted values, measured in response units. Neither is a guarantee about how far any particular prediction will be from the observed response.

The summaries do not replace the graphs. As discussed in “Choosing Graphs to Support a Linear Model” and “Reading a Residual Plot for Model Fit,” examine the scatterplot and residual plot for each candidate. A high \(r^2\) cannot make a curved pattern linear or eliminate a changing spread of residuals. Also consider whether the predictor is available when a prediction is needed and whether its observed range covers the intended use.

A Practical Comparison Process

A clear comparison follows the question outward: establish that the models are comparable, examine their numerical fit, then decide what the evidence means for the stated purpose. There is no automatic rule that says the numerically best-fitting predictor must be the best choice for every situation.

1
Set up a fair comparison.
Confirm that both lines predict the same response and use the same cases. Name the candidate explanatory variables and retain the same response units.
2
Compare \(r^2\) and \(s\).
Report \(r^2\) as a proportion or percentage of response variation accounted for. Report \(s\) in the response variable’s units, and say which model has the larger \(r^2\) and smaller \(s\).
3
Check model behavior and purpose.
Look at each scatterplot and residual plot, and consider the prediction’s intended setting. A modest numerical advantage may not outweigh a clear residual pattern or a predictor that is impractical to measure.
4
Make a qualified choice.
State which model you would use for the stated goal and support the choice with the numerical and contextual evidence. Avoid turning an association into a causal claim.

Worked Example: Comparing Two Predictors of Plant Growth

Worked Example: Fertilizer Amount or Watering Frequency?

Original AP-style question. In an invented greenhouse exercise, five plots have the following fertilizer amounts \(x_1\), watering-frequency values \(x_2\), and plant growth measurements \(y\), in centimeters. Compare the two simple linear models for predicting growth.

PlotFertilizer amount \(x_1\)Watering-frequency value \(x_2\)Growth \(y\) (cm)
11110
22312
33213
44515
55416

Solution. Both models use the same five growth measurements, so \(\bar{y}=13.2\) cm and \(SST=\sum(y-\bar{y})^2=22.8\text{ cm}^2\). For the fertilizer model, the fitted line is \(\hat{y}=8.7+1.5x_1\). Its fitted growth values are \(10.2,11.7,13.2,14.7,16.2\) cm, giving residuals \(-0.2,0.3,-0.2,0.3,-0.2\) cm. Thus:

$$ SSE_1=(-0.2)^2+(0.3)^2+(-0.2)^2+(0.3)^2+(-0.2)^2=0.30 $$

For the watering-frequency model, the fitted line is \(\hat{y}=9.3+1.3x_2\). Its fitted values are \(10.6,13.2,11.9,15.8,14.5\) cm, producing residuals \(-0.6,-1.2,1.1,-0.8,1.5\) cm. Its squared residuals sum to \(0.36+1.44+1.21+0.64+2.25=5.90\). With \(n=5\), calculate both summaries for each model:

$$ r_1^2=1-\frac{0.30}{22.8}\approx 0.9868, \qquad s_1=\sqrt{\frac{0.30}{5-2}}\approx 0.316\text{ cm} $$
$$ r_2^2=1-\frac{5.90}{22.8}\approx 0.7412, \qquad s_2=\sqrt{\frac{5.90}{5-2}}\approx 1.402\text{ cm} $$

Conclusion. For these five plots, the fertilizer model has the larger \(r^2\): about 98.7% compared with 74.1%. Its \(s\) is also smaller: growth measurements are typically about 0.316 cm from its fitted values, compared with about 1.402 cm for the watering-frequency model. These summaries favor fertilizer amount for describing growth in this data set. Before using that line for prediction, inspect its scatterplot and residual plot and consider whether the plots represent the intended setting. The comparison describes association; it does not show that fertilizer caused the observed growth differences.

Worked Example: Compare Models From Summary Information

Worked Example: Predicting Daily Water Use

Original AP-style question. In an invented analysis of the same eight days, a community garden considers average temperature and hours of sunlight as separate predictors of daily water use. Water use is measured in liters. The response has \(SST=140\text{ liters}^2\). The temperature model has \(r^2=0.64\), and the sunlight model has \(r^2=0.49\). Compare the models using \(r^2\) and \(s\).

Solution. The temperature model accounts for 64% of the variation in observed daily water use; the sunlight model accounts for 49%. To compare typical residual sizes, recover each \(SSE\) from \(r^2=1-SSE/SST\), then use \(s=\sqrt{SSE/(n-2)}\). Both models have \(n=8\), so the residual degrees of freedom are \(8-2=6\).

$$ SSE_{\text{temperature}}=(1-0.64)(140)=50.4, \qquad s_{\text{temperature}}=\sqrt{\frac{50.4}{6}} \approx 2.898\text{ liters} $$
$$ SSE_{\text{sunlight}}=(1-0.49)(140)=71.4, \qquad s_{\text{sunlight}}=\sqrt{\frac{71.4}{6}} \approx 3.450\text{ liters} $$

Conclusion. For these same eight days, the temperature model has a higher \(r^2\) and a lower \(s\). It accounts for a larger proportion of the variation in daily water use, and observed use is typically closer to its fitted values by this response-scale measure. If the goal is to describe these observations, temperature is favored by these two summaries. A practical prediction choice should also consider residual plots and whether temperature is known at the time predictions are needed. The summaries alone do not establish a causal effect of temperature.

Worked Example: When the Numerical Advantage Is Small

Worked Example: Comparing Two Recycling Predictors

Original AP-style question. In an invented data set of six collection areas, the same response—recycled material collected per area, in kilograms—has \(SST=50\text{ kg}^2\). A model using number of households has \(r^2=0.81\); a model using collection stops has \(r^2=0.80\). Compare their residual standard deviations and recommend how to decide between them.

Solution. Each model uses six areas, so \(n-2=4\). For households, \(SSE=(1-0.81)(50)=9.5\text{ kg}^2\), and for collection stops, \(SSE=(1-0.80)(50)=10.0\text{ kg}^2\). Therefore:

$$ s_{\text{households}}=\sqrt{\frac{9.5}{4}}\approx 1.541\text{ kg}, \qquad s_{\text{stops}}=\sqrt{\frac{10.0}{4}}\approx 1.581\text{ kg} $$

Conclusion. The households model has a slightly higher \(r^2\) and slightly lower \(s\): it accounts for 81% rather than 80% of the response variation, and its residuals are typically about 1.541 kg rather than 1.581 kg. These summaries give households a small numerical advantage, not an overwhelming one. If collection-stop counts are easier to obtain when planning service, that practical advantage may matter. Check both residual plots and the intended prediction task before recommending either model; do not claim that the tiny difference proves one predictor is meaningfully better in every setting.

Common Mistakes and AP Exam Tips

  • Comparing summaries from different response data. A fair comparison here requires the same response values and cases. Say explicitly that both models predict the same response for the same observations.
  • Reporting \(r^2\) as the percent predicted correctly. Write that the model accounts for a stated percentage of the variation in the response values. It is not a percentage of individual predictions that are correct.
  • Leaving \(s\) without response units. If the response is plant growth in centimeters, \(s\) is in centimeters. Interpret it as a typical residual size, not a guaranteed error limit.
  • Choosing solely by a higher \(r^2\). A larger value favors a model on that summary, but residual plots, prediction purpose, data scope, and predictor practicality also matter. As in “Reliability Within the Range of Data,” a numerical fit summary does not settle whether a particular prediction is dependable.
  • Confusing association with cause and effect. A model comparison describes how variables are associated in the observations. Use “is associated with” unless the study design supports a causal conclusion.

A strong AP response names the two models and the common response, compares \(r^2\) and \(s\) with appropriate units and interpretations, and gives a qualified choice connected to the goal. If the numerical advantage is small, say so rather than presenting it as decisive.

Key takeaway: When two simple linear models with intercepts use the same response values and cases, compare \(r^2\) and \(s\) together: the model with larger \(r^2\) has smaller \(s\). Then check the plots and context before choosing a model for the intended use.

Check Your Understanding

Answer each question using the comparison ideas from this tutorial.

  1. Two models use different predictors but the same response and the same 12 cases. Why can their \(r^2\) and \(s\) be compared directly?
  2. For one model, \(SST=80\) and \(SSE=24\). Find \(r^2\) and interpret it in context, using a response of daily energy use.
  3. For the model in question 2, calculate \(s\) if \(n=8\). Include the response units of kilowatt-hours.
  4. Two models have almost identical \(r^2\) and \(s\). Name two contextual or graphical considerations that could help guide the choice.
  5. Why does a higher \(r^2\) for a predictor not, by itself, show that changing the predictor causes a change in the response?