Tutorials › AP Statistics › Full AP-Style Regression Free-Response Practice

Comparing and communicating regression models · Tutorial 999 of 1000

Full AP-Style Regression Free-Response Practice

Work through a multi-part regression question and practice connecting each claim to the right statistic, calculation, or graph.

Intermediate 10 min read

What You'll Learn

  • Organize a multi-part regression response so every part of the prompt receives a direct answer.
  • Describe a scatterplot’s direction, form, strength, and unusual features in context.
  • Interpret a fitted line, its slope and intercept, \(r^2\), and \(s\) accurately.
  • Calculate and interpret a residual using observed and predicted responses.
  • Use residual evidence and the data’s scope to write a supported conclusion.

Bring the Regression Pieces Together

A complete regression response is more than a collection of correct statistics. It connects the data display, fitted line, residuals, and conclusion to the question being asked. In “Scoring a Model Free-Response Regression Answer,” you practiced checking criteria one at a time. Here, you will work through an original multi-part question and assemble those pieces into one coherent response.

The question below is descriptive: it asks what the observed data show and how well a linear model represents them. As in “Choosing Graphs to Support a Linear Model” and “Selecting Evidence for a Written Conclusion,” use each piece of evidence for the claim it can support. A scatterplot describes the observed association, the regression output summarizes the fitted line, and residuals help assess what the line misses.

Key idea: Answer each part directly, show requested calculations, and keep the final claim within the cases and conditions represented by the data. A correct statistic does not by itself justify a broader claim.

The Practice Question

In an invented observational project, a coach records weekly swim-practice hours and 400-meter swim times for six swimmers on a team. Let \(x\) be practice hours per week and \(y\) be swim time in seconds. The observations are:

Practice hours, \(x\)400-meter time, \(y\) (seconds)
179
276
373
472
568
670

A scatterplot shows a generally straight, downward pattern. There are no clearly separated points, though a small data set makes it difficult to judge fine details. The observed practice-hour values range from 1 to 6. Regression output for predicting swim time from practice hours is:

$$ \hat{y}=80-2x,\qquad r=-0.9354,\qquad r^2=0.875,\qquad s=1.581\text{ seconds}. $$

The question has five parts: (a) describe the association in context; (b) interpret the slope and intercept; (c) interpret \(r^2\) and \(s\); (d) calculate and interpret the residual for the swimmer who practiced 2 hours; and (e) assess the linear model and write a conclusion that respects the data’s scope.

Response plan: First describe the observed pattern. Then interpret the numerical output with context and units. Calculate the requested residual as observed response minus fitted response. Finally, use the residual evidence and the way the data were collected to state a appropriately limited conclusion.

Worked Example: A Complete Multi-Part Regression Response

Worked Example: Weekly Swim Practice and 400-Meter Time

State. The goal is to describe the association between weekly practice hours and 400-meter time for these six swimmers, interpret the fitted line and summary statistics, check a residual, and evaluate the conclusion the data support.

Plan. Use the scatterplot to describe direction, form, and strength, and note any unusual features. Use the regression output for the line’s interpretations. For the residual, substitute the swimmer’s practice hours into the fitted line and subtract the predicted time from the observed time. Use the residual information and observational nature of the project to limit the conclusion. As in “Reading a Residual Plot for Model Fit,” a model’s fit is not judged by one statistic alone.

Do: describe the association. The scatterplot shows a strong, negative, approximately linear association between weekly practice hours and 400-meter swim time among the six swimmers. In general, swimmers with more weekly practice hours had lower times. There are no clearly separated outliers in this small set. The word “strong” is supported by the fairly close, straight overall pattern and the large magnitude of \(r\), not by \(r^2\) being an “accuracy” measure.

Do: interpret the line. The slope is \(-2\) seconds per weekly practice hour. In context, for these swimmers, each additional hour of weekly practice is associated with a predicted decrease of 2 seconds in 400-meter time according to the fitted line. This describes a change in predicted response, not a guaranteed change for every swimmer.

The intercept is 80 seconds. It is the model’s predicted 400-meter time for a swimmer who practices 0 hours per week. Since the observed practice-hour values range from 1 to 6, 0 hours is outside the observed range. The intercept has a mathematical interpretation, but it may not be useful for describing these swimmers.

Do: interpret \(r^2\) and \(s\). The coefficient of determination is \(r^2=0.875\), or \(87.5\%\). Thus, 87.5% of the variation in 400-meter swim times among these six swimmers is accounted for by the linear regression of swim time on weekly practice hours. This does not mean that practice causes 87.5% of a swimmer’s time or that predictions are 87.5% accurate.

The value \(s=1.581\) seconds describes the typical size of the residuals from this fitted line for these swimmers. In context, actual swim times typically differ from their fitted times by about 1.581 seconds. This is a typical size, not a guarantee that every swimmer’s residual is within 1.581 seconds.

Do: calculate and interpret the residual. For the swimmer who practiced 2 hours, the fitted time is:

$$ \hat{y}=80-2(2)=76\text{ seconds}. $$

The observed time is 76 seconds, so the residual is:

$$ \text{residual}=y-\hat{y}=76-76=0\text{ seconds}. $$

This swimmer’s observed time is exactly the value predicted by the fitted line. A residual of zero describes this observation; it does not mean that the model predicts every swimmer perfectly.

Do: assess the model and its scope. The scatterplot is approximately linear, and the residuals do not show a clear curved pattern. The residual sizes are not identical, and there are only six swimmers, so this is limited evidence for the model’s fit rather than proof that a line captures every feature of the relationship. The project is observational, so it supports an association, not a claim that additional practice causes faster times. The data represent these six swimmers; without information about how they were selected, the result should not automatically be generalized to other swimmers.

Conclude. Among the six swimmers studied, more weekly practice hours are associated with lower 400-meter swim times in a strong, approximately linear pattern. The fitted line predicts a 2-second decrease in time per additional weekly practice hour, and it accounts for 87.5% of the observed variation in times. The residual evidence does not show a clear curvature, but the small observational data set does not establish causation or justify broad generalization.

Check the Residual and Fit Separately

The multi-part response used a zero residual, which can make the calculation look trivial. A nonzero residual is just as important to interpret carefully: its sign indicates whether the observed response is above or below the fitted value. The size is measured in the response’s units.

Worked Example: A Nonzero Residual for a Water-Use Model

Situation. In an invented analysis of garden plots, let \(x\) be watering time in minutes and \(y\) be water used in liters. A fitted line is \(\hat{y}=10+1.55x\). One plot was watered for 12 minutes and used 31 liters. The observed watering times ranged from 4 to 16 minutes.

Calculate the fitted value. Substitute \(x=12\) into the line:

$$ \hat{y}=10+1.55(12)=10+18.6=28.6\text{ liters}. $$

Calculate the residual. Subtract the fitted value from the observed water use:

$$ \text{residual}=y-\hat{y}=31-28.6=2.4\text{ liters}. $$

Interpret and assess. This plot used 2.4 liters more water than the model predicted for a 12-minute watering time, so its point lies above the fitted line. Twelve minutes is within the observed range of 4 to 16 minutes, so this calculation is not an extrapolation. The residual for this one plot does not by itself establish whether the model fits well overall; that judgment should consider the overall scatterplot and residual plot.

This example reinforces a distinction made in “Residual Plot Versus Scatterplot Diagnosis”: a single residual tells you how far one observation is vertically from its fitted value, while a residual plot can reveal systematic patterns across many observations. A small residual for one case is not enough to establish that a model is appropriate.

Connect the Numbers to a Conclusion

A multi-part question may ask for several interpretations before it asks for a conclusion. Do not simply repeat every number. Select the evidence that supports the conclusion, explain what it means in context, and avoid claims that the study design cannot support. The slope describes predicted change; \(r\) describes direction and strength of linear association; \(r^2\) describes the proportion of response variation accounted for; and \(s\) describes typical residual size.

Worked Example: Revise an Overstated Regression Conclusion

Situation. In an invented observational project, students record daily recreational screen time, \(x\), in hours and nightly sleep, \(y\), in hours for 20 students at one school. The observed screen-time values range from 2 to 8 hours. The fitted line is \(\hat{y}=9.2-0.35x\), with \(r=-0.72\), \(r^2=0.5184\), and \(s=0.8\) hours. A student writes: “Screen time causes students everywhere to lose 0.35 hours of sleep for every extra hour on a device. The model is 51.84% accurate.”

Identify the problems. The slope is being treated as a causal effect and as a guaranteed change. The data are observational, so the analysis establishes an association, not that screen time causes sleep to decrease. The conclusion also generalizes from students at one school to “students everywhere” without support. Finally, \(r^2\) is not a measure of accuracy.

Revise with the output’s correct meanings. The slope says that, among the students studied, each additional hour of daily screen time is associated with a predicted decrease of 0.35 hours of nightly sleep according to the fitted line. The \(r^2\) interpretation is:

$$ 0.5184\times100\%=51.84\%. $$

So, 51.84% of the variation in nightly sleep among these 20 students is accounted for by the linear regression of sleep on screen time. The value \(s=0.8\) hours means that residuals typically have a size of about 0.8 hours for these students. A careful conclusion is: “Among the 20 students at this school, greater daily screen time is associated with less nightly sleep in the fitted linear model. These observational data do not show that screen time causes less sleep, and the results should not automatically be generalized to students elsewhere.”

The revised conclusion does not discard the evidence; it states what the evidence can support. This is the same discipline practiced in “Writing a Conclusion Without Overclaiming” and “Describing Limitations of a Regression Analysis.”

Common Mistakes and AP Exam Tips

  • Listing statistics without describing the pattern. If asked to describe an association, state direction, form, and strength in context. A value of \(r\) alone does not fully answer that request.
  • Turning association into cause. Unless the data come from an appropriate randomized experiment, use wording such as “is associated with” rather than “causes.”
  • Interpreting the slope without units. Include predicted response change per one explanatory-variable unit. For example, name both seconds and weekly practice hours.
  • Giving the intercept too much importance. Interpret it as the predicted response at \(x=0\), then check whether zero is meaningful and within the observed \(x\)-range.
  • Calling \(r^2\) accuracy. A full-credit interpretation identifies the response variation, the cases, and the linear model. It does not claim that the model predicts that percentage of individual responses correctly.
  • Reversing the residual subtraction. The residual is observed response minus fitted response, \(y-\hat{y}\). A positive residual is above the line; a negative residual is below it.
  • Judging the whole model from one residual. Use a residual plot to look for systematic structure, and describe the evidence rather than claiming a model is perfect.
  • Making the conclusion broader than the data. Name the cases or setting represented and note meaningful limits, such as a small sample or observational data.
  • Leaving a part unanswered while polishing another. Read the lettered parts again before finishing. If the prompt asks for a residual, a fit assessment, and a conclusion, make each one visible in the response.

For a strong response, a reader should be able to locate the description, interpretations, calculation, model assessment, and conclusion without guessing what you meant. You do not need lengthy prose for every part, but each statement should include the context and statistical meaning required by the question.

Key takeaway: A complete regression free response connects the observed pattern, fitted model, residual evidence, and conclusion. Match each claim to the evidence that supports it, show requested calculations, and state the limits of what the data establish.

Check Your Understanding

Use the swim-practice example and the ideas in this tutorial to answer each question.

  1. Write one sentence describing the direction, form, and strength of the association between weekly practice hours and 400-meter time.
  2. For the fitted line \(\hat{y}=80-2x\), what does the slope mean in context, and why should it not be phrased as a guaranteed change?
  3. For the swimmer who practiced 5 hours and had an observed time of 68 seconds, calculate the fitted time and residual. Interpret the sign of the residual.
  4. What does \(r^2=0.875\) describe in this setting? Name one interpretation it does not support.
  5. Give one reason the project cannot establish that additional practice causes faster swim times, and one reason the results may not apply to all swimmers.