Tutorials › AP Statistics › Linear Model Interpretation Mixed Review

Linear regression models · Tutorial 860 of 1000

Linear Model Interpretation Mixed Review

Bring regression equations and summary statistics together: interpret coefficients, calculate a line, and check predictions in context.

Intermediate 9 min read

What You'll Learn

  • Identify the predictor and response before interpreting a regression equation.
  • Interpret a slope as a change in predicted response per predictor unit, with units.
  • Interpret an intercept at predictor value zero and assess whether that value is supported by the data.
  • Calculate a least-squares line from \(r\), sample means, and sample standard deviations.
  • Check a calculated line by verifying its mean point and comparing predicted differences.
  • Distinguish predictions from actual outcomes and association from causation.

Bring the Pieces of a Regression Line Together

In “Free-Response Practice Interpreting a Regression Line,” you practiced writing clear interpretations of slopes and intercepts. This mixed review combines those skills with calculations from summary statistics. You will read a supplied equation, build an equation from \(r\), sample means, and sample standard deviations, and check whether the result makes sense in context.

The main challenge is keeping the roles straight while moving between numbers and meaning. The predictor \(x\) is used to predict the response \(y\), and the fitted line gives a predicted response \(\hat{y}\). As in “Reading the Equation of a Regression Line,” an equation written as \(\hat{y}=a+bx\) uses \(a\) for the intercept and \(b\) for the slope.

Mixed-review strategy: First identify the predictor, response, and units. Then interpret or calculate the coefficients. Finally, check the line with a prediction, the observed predictor range, or the fact that a least-squares line with an intercept passes through \((\bar{x},\bar{y})\).

Read an Equation, Then Check Its Predictions

When an equation is supplied, do not start by calculating predictions. First translate its coefficients into context. The slope describes the change in the line’s predicted response for a one-unit increase in the predictor. Its units are response units per predictor unit. The intercept is the predicted response at \(x=0\); whether that prediction is practically useful is a separate question.

A useful arithmetic check is to compare predictions at two predictor values. Their difference should equal the slope multiplied by the difference between those predictor values. This connects the equation’s numerical behavior to the slope interpretation. It does not say that actual responses must differ by exactly that amount.

Worked Example: Check a Speaker-Battery Model

A fictional electronics shop records hours of audio playback, \(x\), and the battery remaining, \(y\), as a percentage. For the observed speakers, playback time ranges from 2 to 9 hours. A regression equation is:

$$ \widehat{\text{battery remaining}}=96-4.2(\text{playback hours}) $$

Interpret the slope and intercept, then find and check the predicted difference in battery remaining between 5 and 8 hours of playback.

The slope is \(-4.2\), measured in percentage points per playback hour. Thus, for each additional hour of playback, the regression line predicts a decrease of 4.2 percentage points in battery remaining. The intercept is 96 percentage points: the line predicts 96% battery remaining at zero hours of playback. Zero is below the observed range of 2 to 9 hours, so this intercept prediction is an extrapolation and may not be well supported by these observations.

Substitute 5 and 8 into the equation:

$$ \hat{y}(5)=96-4.2(5)=75 \qquad \hat{y}(8)=96-4.2(8)=62.4 $$

The predicted difference, comparing 8 hours with 5 hours, is \(62.4-75=-12.6\) percentage points. Check this using the slope and predictor difference: \(-4.2(8-5)=-4.2(3)=-12.6\) percentage points. The two calculations agree. In context, the line predicts 12.6 percentage points less battery remaining at 8 hours than at 5 hours. This is a difference between predictions, not a guarantee about any particular speaker.

Calculate a Line from Summary Statistics

Sometimes a question gives a correlation and summary statistics instead of a regression equation. For a least-squares regression line with an intercept, use the relationship from “Computing Slope From \(r\) and Summary Statistics” to calculate the slope, then use the sample means to calculate the intercept.

$$ b=r\left(\frac{s_y}{s_x}\right) \qquad a=\bar{y}-b\bar{x} \qquad \hat{y}=a+bx $$

Here, \(s_x\) and \(s_y\) are the sample standard deviations of the predictor and response. The ratio \(s_y/s_x\) converts the unitless correlation into slope units. After finding the line, interpret its coefficients in context. The line should also pass through \((\bar{x},\bar{y})\), as explained in “Finding the Line Through the Means.”

Worked Example: Build a Free-Throw Practice Model

A fictional basketball program records each player’s weekly practice time, \(x\), in hours, and the player’s successful free throws, \(y\), as a percentage during a practice assessment. The summary statistics are \(\bar{x}=4.5\) hours, \(\bar{y}=68\%\), \(s_x=1.2\) hours, \(s_y=10\) percentage points, and \(r=0.60\). The observed practice times range from 2.5 to 6.5 hours. Find the regression line, interpret its coefficients, and predict the assessment percentage for a player who practices 6 hours per week.

State. The predictor is weekly practice time in hours; the response is the successful-free-throw percentage. We need the least-squares line and its prediction at 6 hours.

Plan. Calculate the slope using \(b=r(s_y/s_x)\), then calculate the intercept using \(a=\bar{y}-b\bar{x}\). Keep the units attached: the slope is percentage points per hour, and the intercept is percentage points.

Do. The slope calculation is:

$$ b=0.60\left(\frac{10}{1.2}\right) =0.60(8.333\ldots) =5 $$

As a check, \(5(1.2)=6\), and \(0.60(10)=6\), so the value of \(b\) satisfies \(b s_x=r s_y\). Next, calculate the intercept:

$$ a=68-5(4.5)=68-22.5=45.5 $$

The fitted line is:

$$ \widehat{\text{free-throw percentage}} =45.5+5(\text{practice hours}) $$

At 6 hours, the predicted percentage is:

$$ \hat{y}=45.5+5(6)=45.5+30=75.5\% $$

Conclude. For each additional hour of weekly practice, the regression line predicts a 5-percentage-point increase in the successful-free-throw percentage. The intercept says the line predicts 45.5% for a player with zero practice hours per week. Since zero is outside the observed range of 2.5 to 6.5 hours, that intercept is an extrapolation and may not be a useful practical prediction. For a player who practices 6 hours per week, which is within the observed range, the predicted percentage is 75.5%.

Finally, check the mean point. Substituting the mean practice time gives \(45.5+5(4.5)=45.5+22.5=68\%\), equal to \(\bar{y}\). The calculated line therefore passes through \((4.5,68)\), as it should. This check can reveal an arithmetic or substitution error, although passing it alone does not prove that every other calculation is correct.

Use the Sign and Context to Review a Negative Slope

A negative correlation produces a negative slope when both variables vary, because the standard-deviation ratio in \(b=r(s_y/s_x)\) is positive. The predicted response therefore decreases as the predictor increases. The units and context still matter: the sign tells you the direction, while the variable names and units explain what is changing.

A regression line summarizes an observed linear pattern. As discussed in “Regression Does Not Mean Causation in Slope Statements,” interpreting a negative slope does not establish that increasing the predictor causes the response to decrease. The interpretation should describe the line’s prediction and avoid promising an exact outcome for an individual case.

Worked Example: Model Irrigation Need from Rainfall

A fictional greenhouse team records weekly rainfall, \(x\), in millimeters, and the amount of irrigation water used, \(y\), in liters. The summary statistics are \(\bar{x}=22\) millimeters, \(\bar{y}=40\) liters, \(s_x=8\) millimeters, \(s_y=12\) liters, and \(r=-0.50\). Rainfall values range from 8 to 36 millimeters. Find the regression line, interpret the slope, and predict irrigation use for a week with 28 millimeters of rain.

First calculate the slope:

$$ b=-0.50\left(\frac{12}{8}\right) =-0.50(1.5) =-0.75\text{ liters per millimeter} $$

Check the result against the slope relationship: \((-0.75)(8)=-6\), and \((-0.50)(12)=-6\). Now find the intercept:

$$ a=40-(-0.75)(22) =40+16.5 =56.5\text{ liters} $$

The line is:

$$ \widehat{\text{irrigation use}} =56.5-0.75(\text{rainfall}) $$

For each additional millimeter of weekly rainfall, the regression line predicts 0.75 fewer liters of irrigation water used. This is an association in the observed data, not proof that rainfall alone caused irrigation use to change. The intercept says the line predicts 56.5 liters of irrigation use at zero millimeters of rainfall. Since zero is outside the observed range of 8 to 36 millimeters, that prediction is an extrapolation.

For 28 millimeters of rain, direct substitution gives:

$$ \hat{y}=56.5-0.75(28) =56.5-21 =35.5\text{ liters} $$

A second check uses the mean point and the change from the mean. The mean point is \((22,40)\); moving from 22 to 28 millimeters is an increase of 6 millimeters. The predicted response changes by \((-0.75)(6)=-4.5\) liters, so \(40-4.5=35.5\) liters. This agrees with direct substitution. Because 28 millimeters is within the observed rainfall range, this prediction is an interpolation rather than an extrapolation.

A Compact Review Workflow

For a mixed question, separate the calculation from the interpretation. Use this sequence whether the equation is given or must be built from summary statistics.

1
Set the roles.
Identify the predictor, response, and units. Check that the equation predicts the response from the predictor.
2
Find or read the coefficients.
If needed, calculate \(b=r(s_y/s_x)\) and \(a=\bar{y}-b\bar{x}\). Keep the sign and units of each coefficient.
3
Translate the coefficients.
Describe the slope as a predicted response change for a one-unit predictor increase. Describe the intercept as the predicted response when the predictor is zero.
4
Check the result.
Verify the mean point for a line calculated from summary statistics. For a prediction difference, compare direct substitutions with slope times the predictor change.
5
Put the prediction in context.
Compare the predictor value with the observed range, and distinguish an interpolation from an extrapolation when relevant. Describe predictions, not guaranteed actual outcomes.

Common Mistakes and AP Exam Tips

  • Reversing the variables. The slope is response units per predictor unit, not predictor units per response unit. State what \(x\) measures and what \(y\) measures before interpreting the coefficient.
  • Calling a slope a change in actual outcomes. Say “the regression line predicts” a change in the response. Do not imply that every individual response changes by exactly the slope amount.
  • Dropping the units. A slope of \(-0.75\) is not complete by itself. In the rainfall example, it is \(-0.75\) liters per millimeter, so each additional millimeter corresponds to a predicted decrease of 0.75 liters.
  • Misreading the intercept. An intercept is not automatically an average or typical response. State the predicted response specifically at predictor value zero, then separately explain if zero is outside the observed range.
  • Making a sign error when finding the intercept. In \(a=\bar{y}-b\bar{x}\), a negative slope means subtracting a negative quantity. Check the result by confirming that \(a+b\bar{x}=\bar{y}\).
  • Treating a check as proof of everything. A mean-point check can catch some arithmetic mistakes, but it does not replace checking the predictor and response, units, coefficient signs, and prediction range.
  • Using causal language without support. A negative slope describes a downward predicted pattern. It does not, by itself, show that changing the predictor causes a change in the response.

A full-credit style response makes each link visible: the predictor, the predicted response, the direction, and the units. When calculating a line, show the slope and intercept formulas with their substitutions; when interpreting it, connect each coefficient to the situation. Keep practical concerns—such as an intercept outside the observed range—separate from the coefficient’s mathematical meaning.

Key takeaway: Read or calculate the line with the predictor and response roles fixed. Interpret the slope as predicted response change per predictor unit and the intercept as the predicted response at zero. Check calculations with the mean point or a prediction difference, and judge predictions in context.

Check Your Understanding

For each item, show the relevant calculation and use the variables’ context when interpreting a result.

  1. A model predicts package weight in grams from the number of items packed, with equation \(\hat{y}=120+35x\). Interpret the slope, including its units.
  2. For a line \(\hat{y}=18-2.5x\), find the predicted response at \(x=4\) and at \(x=7\). Check the predicted difference using the slope.
  3. Given \(r=0.40\), \(s_x=3\), and \(s_y=15\), calculate the regression slope. Include its units if \(x\) is measured in minutes and \(y\) in points.
  4. A least-squares line has slope \(b=4\), \(\bar{x}=5\), and \(\bar{y}=31\). Calculate the intercept and check that the line passes through the mean point.
  5. A line predicts water use from rainfall, but the observed rainfall ranges from 10 to 40 millimeters. What should an answer mention when discussing the practical usefulness of the intercept at zero rainfall?