Tutorials › AP Statistics › Unit 5 Regression Synthesis Review

Comparing and communicating regression models · Tutorial 1000 of 1000

Unit 5 Regression Synthesis Review

Practice combining regression evidence into clear, context-aware conclusions without overstating what a model shows.

Intermediate 10 min read

What You'll Learn

  • Connect a scatterplot’s direction, form, and strength to correlation and a fitted regression line.
  • Calculate a least-squares line from summary values and check its predictions with residuals.
  • Interpret \(r^2\) and \(s\) as different summaries of model fit, with the correct context and units.
  • Compare two candidate predictors when models use the same response and cases.
  • Identify when extrapolation, observational data, or a small data set limits a regression conclusion.

Regression as a Connected Argument

A strong regression response connects several kinds of evidence rather than treating each statistic as a separate answer. The scatterplot shows the observed relationship; the correlation and fitted line summarize its linear pattern; residuals show how observed values differ from fitted values; and \(r^2\) and \(s\) describe different aspects of fit. The conclusion then states what these results support—and what they do not.

In “Full AP-Style Regression Free-Response Practice,” you assembled these parts for one analysis. This final synthesis reviews them in a compact practice set. The new technique is to use an evidence chain: for each claim, identify the graph, statistic, or data feature that directly supports it, then check whether the study’s scope permits the claim.

Key idea: A good regression conclusion is not a list of numbers. It links the observed pattern, the model’s summaries, and the data’s limitations to the question being asked.

Practice Set: Heating Use and Outdoor Temperature

In an invented monitoring project, daily outdoor temperature and household heating use were recorded on six winter days. Let \(x\) be the daily outdoor temperature in degrees Celsius and \(y\) be heating energy use in kilowatt-hours (kWh). The data are:

Temperature, \(x\) (°C)Heating use, \(y\) (kWh)
114
212
311
49
58
67

The scatterplot has a strong, downward, approximately linear pattern. The temperatures range from 1°C to 6°C. The following calculations give the least-squares regression line and its summaries. Here \(S_{xx}\), \(S_{yy}\), and \(S_{xy}\) denote the sums of squared deviations in \(x\), squared deviations in \(y\), and cross-products of deviations, respectively.

Worked Example: Build and Interpret a Regression Summary

Worked Example: Heating Use on Six Winter Days

State. Describe the association between outdoor temperature and heating use on these six days, calculate the fitted line and its main summaries, and assess what conclusion the data support.

Plan. Use the scatterplot for direction, form, and strength. Calculate the least-squares slope and intercept from the summary values, then use residuals to check what the fitted line misses. Interpret \(r^2\) and \(s\) in context. Finally, consider the small number of days and the observational nature of the records before making a broader claim.

Do: calculate the line and summaries. The means are \(\bar{x}=3.5\) degrees Celsius and \(\bar{y}=61/6=10.1667\) kWh. From the data, \(\sum x=21\), \(\sum x^2=91\), \(\sum y=61\), \(\sum y^2=655\), and \(\sum xy=189\). Therefore:

$$ S_{xx}=91-\frac{21^2}{6}=17.5,\qquad S_{yy}=655-\frac{61^2}{6}=\frac{209}{6}=34.8333, $$
$$ S_{xy}=189-\frac{(21)(61)}{6}=-24.5. $$

The slope is \(b=S_{xy}/S_{xx}\), and the intercept is \(a=\bar{y}-b\bar{x}\):

$$ b=\frac{-24.5}{17.5}=-1.4,\qquad a=10.1667-(-1.4)(3.5)=15.0667. $$

Thus, the fitted line, with coefficients rounded to three decimal places, is:

$$ \hat{y}=15.067-1.400x. $$

The slope means that on these days, each 1°C increase in outdoor temperature is associated with a predicted decrease of 1.4 kWh in heating use, according to this line. This is an average change in predicted response, not a guaranteed change on every day.

The correlation is calculated from the cross-product and sums of squares:

$$ r=\frac{S_{xy}}{\sqrt{S_{xx}S_{yy}}} =\frac{-24.5}{\sqrt{(17.5)(34.8333)}}\approx -0.9923. $$

The negative sign indicates a negative linear association, and the magnitude indicates a very strong linear association in these six observations. Squaring the correlation gives:

$$ r^2=(-0.9923)^2\approx 0.9847. $$

About 98.47% of the variation in heating use among these six days is accounted for by the linear regression of heating use on outdoor temperature. This is not a claim that the model is 98.47% accurate or that temperature causes that percentage of heating use.

Do: check residuals and typical residual size. For the day with \(x=4\)°C, the fitted value is:

$$ \hat{y}=15.0667-1.4(4)=9.4667\text{ kWh}. $$

The observed use that day was 9 kWh, so its residual is:

$$ y-\hat{y}=9-9.4667=-0.4667\text{ kWh}. $$

The negative residual means heating use was about 0.467 kWh below the fitted value for that day. Across all six observations, the residual sum of squares is \(SSE=S_{yy}-bS_{xy}=34.8333-(-1.4)(-24.5)=0.5333\). The residual standard deviation is:

$$ s=\sqrt{\frac{SSE}{n-2}} =\sqrt{\frac{0.5333}{6-2}} \approx 0.365\text{ kWh}. $$

Thus, residuals from this line typically have a size of about 0.365 kWh for these days. As discussed in “Reading a Residual Plot for Model Fit,” the residual plot—not \(r\) or \(r^2\) alone—helps reveal systematic patterns the line leaves behind. These six residuals do not show an obvious sustained curve, but six observations provide limited evidence about model fit.

Conclude. On these six winter days, higher outdoor temperatures were associated with lower heating use in a strong, approximately linear pattern. The fitted line accounts for about 98.47% of the observed variation in heating use, with typical residual size about 0.365 kWh. These records describe an association over six days; they do not establish that temperature alone caused the differences in heating use or show that the same line applies in other conditions.

Compare Models Using the Right Evidence

A regression synthesis question may offer two candidate explanatory variables for the same response. As in “Comparing Models With Different Explanatory Variables,” a fair descriptive comparison requires the same response variable, cases, and response scale. When those requirements hold, compare \(r^2\) and \(s\) together: a larger \(r^2\) means a larger proportion of response variation is accounted for, while a smaller \(s\) means smaller typical residuals in response units. For the same response values and cases, these summaries agree in direction: larger \(r^2\) goes with smaller \(s\), apart from rounding.

Worked Example: Choose Between Two Candidate Predictors

Situation. An analyst considers two predictors for heating use on the same six days: outdoor temperature and wind speed. Both regressions use the same six heating-use values in kWh. Suppose the software reports the following rounded summaries:

Explanatory variable\(r\)\(r^2\)\(s\) (kWh)
Outdoor temperature-0.9000.81001.286
Wind speed-0.9500.90250.921

Compare the summaries. Wind speed has the larger absolute correlation, \(0.950\) compared with \(0.900\), and the larger \(r^2\), \(0.9025\) compared with \(0.8100\). It accounts for a larger proportion of heating-use variation in these six days: 90.25% rather than 81.00%. Its \(s\) is also smaller, so its residuals are typically smaller in kWh for these cases.

Check that the summaries fit together. The response variation is \(S_{yy}=34.8333\) and \(n=6\), as calculated in the first example. For outdoor temperature, the residual standard deviation implied by \(r^2=0.8100\) is:

$$ s=\sqrt{\frac{(1-r^2)S_{yy}}{n-2}} =\sqrt{\frac{(1-0.8100)(34.8333)}{4}} \approx 1.286\text{ kWh}. $$

For wind speed, the corresponding calculation is:

$$ s=\sqrt{\frac{(1-0.9025)(34.8333)}{4}} \approx 0.921\text{ kWh}. $$

Conclude cautiously. Considering these descriptive summaries, wind speed provides the better linear fit to heating use for these six days. That does not establish that wind speed is the better predictor in other settings, that its association is causal, or that the linear model is appropriate without checking its scatterplot and residual plot. The candidate-model comparison identifies a better fit among the two summaries, not a complete explanation of heating use.

This comparison also illustrates why statistics must be matched to the claim. \(r\) gives direction and strength of linear association; \(r^2\) describes the proportion of response variation accounted for; and \(s\) describes typical residual size in response units. As noted in “Selecting Evidence for a Written Conclusion,” no single number answers every question about a regression.

Check Predictions and Limits Before Concluding

A fitted line can produce a numerical prediction even when the input value is outside the observed range. That calculation does not make the prediction dependable. “Prediction Limits Mixed Practice Set” emphasizes checking the \(x\)-range separately from whether the predicted response is possible or reasonable. The next example applies both checks and also considers the study design.

Worked Example: Audit a Prediction and Its Claim

Situation. In an invented observational project, a technician records sunlight hours \(x\) and battery charging time \(y\), in minutes, for outdoor charging sessions. The observed sunlight values range from 4 to 12 hours. The fitted line is \(\hat{y}=42-2.05x\). For one session with 10 hours of sunlight, the observed charging time is 23 minutes. A report predicts charging time at 14 hours and claims: “More sunlight causes battery charging time to drop by exactly 2.05 minutes per hour for all outdoor sessions.”

Calculate and interpret the observed residual. For \(x=10\), the fitted charging time is:

$$ \hat{y}=42-2.05(10)=21.5\text{ minutes}. $$

The residual is observed minus fitted:

$$ y-\hat{y}=23-21.5=1.5\text{ minutes}. $$

This session took 1.5 minutes longer to charge than the line predicted. The input, 10 hours, is within the observed range, but this single residual does not establish the overall fit; the scatterplot and residual plot are needed to assess the pattern across sessions.

Check the prediction separately. At 14 hours, the model gives:

$$ \hat{y}=42-2.05(14)=13.3\text{ minutes}. $$

A charging time of 13.3 minutes is possible, but 14 hours is beyond the observed range of 4 to 12 hours. The value is therefore an extrapolation, and the data provide limited support for relying on it. Possibility does not make an out-of-range prediction well supported.

Revise the claim. The slope describes an estimated decrease of 2.05 minutes in predicted charging time per additional hour of sunlight, according to this fitted line. It is not an exact change for every session. Because the data are observational, the association does not by itself show that more sunlight causes shorter charging time; other conditions could differ across sessions. A careful report would describe the association among the sessions studied, identify the observed sunlight range, and treat the 14-hour prediction as uncertain extrapolation.

Common Mistakes and AP Exam Tips

  • Giving a number without its role. Explain what statistic answers the question. For example, \(r^2\) describes response variation accounted for by the linear model; \(s\) describes typical residual size in response units.
  • Using \(r\) to claim model fit is fully checked. A strong correlation does not rule out curvature or changing spread. Refer to the scatterplot and residual plot when assessing whether a line is appropriate.
  • Confusing residual direction. Calculate observed minus fitted, \(y-\hat{y}\). A positive residual is above the line, and a negative residual is below it.
  • Comparing models that are not comparable. Before comparing \(r^2\) and \(s\), verify that both models use the same response variable, response scale, and cases.
  • Calling \(r^2\) “accuracy.” A full-credit interpretation names the response variation, the cases, and the linear regression. It does not describe the percentage of individual predictions that are correct.
  • Overlooking the observed \(x\)-range. State whether a prediction is inside or outside the observed range, then separately consider whether the predicted response is possible or reasonable.
  • Claiming cause from an observational pattern. Use “is associated with” unless the study design justifies a causal conclusion. A fitted slope alone does not establish cause.
  • Generalizing beyond the cases without support. Identify the observations represented and mention limits such as a small data set, a narrow range, or an unclear selection process when relevant.

A useful final check is to match each sentence in your conclusion to evidence: the scatterplot supports the description of the observed pattern; regression output supports interpretations of the line and summaries; residuals and residual plots show what the line misses; and information about the data collection limits the scope of the claim.

Key takeaway: Build a regression conclusion from connected evidence: describe the observed pattern, interpret the fitted line and summaries correctly, check residuals and prediction range, and keep the claim within what the data and study design support.

Check Your Understanding

Use the heating-use data and the worked examples to answer each question.

  1. For the fitted line \(\hat{y}=15.067-1.400x\), interpret the slope in context, including units.
  2. For the day with \(x=5\)°C and observed heating use of 8 kWh, calculate the fitted value and residual. Explain what the residual’s sign means.
  3. Interpret \(r^2\approx0.9847\) in context and state one conclusion it does not support.
  4. When comparing two models for the same response and cases, what do \(r^2\) and \(s\) each tell you? Why should both be considered?
  5. For the battery model, explain why the 14-hour prediction needs caution even though 13.3 minutes is a possible charging time. Give one reason the report’s causal wording is unsupported.