Tutorials › AP Statistics › Choosing Numbers to Support a Linear Model

Comparing and communicating regression models · Tutorial 983 of 1000

Choosing Numbers to Support a Linear Model

Choose the regression numbers that answer the question, and explain what each one contributes in context.

Intermediate 10 min read

What You'll Learn

  • Match \(r\), \(r^2\), \(s\), slope, and intercept to the questions they answer.
  • Interpret correlation’s direction and strength without treating it as proof of causation.
  • Explain \(r^2\) as the proportion of response variation accounted for by the linear model.
  • State the slope and residual standard deviation with appropriate units and context.
  • Decide when an intercept is meaningful and when it is only a feature of the fitted equation.
  • Combine numerical evidence with graphs to support a qualified conclusion about a linear model.

Different Numbers Answer Different Questions

As discussed in “Choosing Graphs to Support a Linear Model,” a scatterplot shows the observed relationship and a residual plot helps check what a fitted line misses. Numerical summaries add specific information: how strong and which direction the linear association runs, how much response variation the model accounts for, how large residuals typically are, and what the fitted line predicts.

Those numbers are not interchangeable. A strong correlation does not tell you how many response units the response changes per unit of \(x\). A large \(r^2\) does not tell you the typical size of a residual in response units. And a fitted intercept is not automatically a useful prediction. Choose evidence that answers the question being asked, and interpret it in context.

Definition: In the least-squares regression line \(\hat{y}=a+bx\), \(r\) is the correlation between the observed \(x\)- and \(y\)-values, \(r^2\) is the coefficient of determination, \(s\) is the residual standard deviation, \(b\) is the slope, and \(a\) is the intercept.

Each number has a particular role. The correlation \(r\) summarizes the direction and strength of the linear association, and has no units. The coefficient of determination \(r^2\) describes the proportion of variation in the observed response values accounted for by the linear model; it also has no units and does not indicate direction. The residual standard deviation \(s\) describes the typical vertical distance of observed responses from the fitted line, in response-variable units.

The slope \(b\) gives the change in the predicted response for a one-unit increase in the explanatory variable, on average according to the fitted line. Its units are response units per explanatory-variable unit. The intercept \(a\) is the predicted response when \(x=0\), measured in response units. Whether that prediction is meaningful depends on the context and the observed range of \(x\).

What each number adds:
  • \(r\): direction and strength of the linear association.
  • \(r^2\): the fraction, or percentage, of response variation accounted for by the linear model.
  • \(s\): a typical residual size in response units.
  • Slope \(b\): the model’s predicted response change for a one-unit increase in \(x\).
  • Intercept \(a\): the model’s predicted response at \(x=0\), if that value has a sensible interpretation.

Match the Evidence to the Question

When asked to describe direction and strength, report \(r\) and its sign. When asked how much variation the model accounts for, report \(r^2\), preferably as a percentage and in context. When the question concerns prediction accuracy or typical vertical scatter, report \(s\) with response units. When describing how the fitted response changes as \(x\) changes, interpret the slope. Discuss the intercept when the predicted response at \(x=0\) matters, and explain its limitations if zero is outside the observed range or does not make sense.

A useful answer often combines more than one number. For example, \(r^2\) gives a proportion of variation, while \(s\) gives a typical error size on the response scale. Neither replaces the other. Also, these summaries describe the fitted linear relationship; they do not establish that changing \(x\) causes a change in \(y\). Consider them alongside the plots and the context, rather than using one impressive-looking statistic as the whole argument.

Worked Example: Support a Model of Practice Time and Score

Worked Example: Practice Time and a Quiz Score

Original AP-style question. In an invented set of five observations, \(x\) is minutes spent practicing and \(y\) is a quiz score in points. The pairs are \((1,52),(2,58),(3,61),(4,69),(5,70)\). Find and interpret \(r\), \(r^2\), \(s\), the slope, and the intercept as evidence about the fitted line.

Solution. The means are \(\bar{x}=3\) minutes and \(\bar{y}=62\) points. The sum of squared deviations in \(x\) is \(S_{xx}=10\), and the sum of cross-products is \(S_{xy}=47\). Thus the slope is \(b=S_{xy}/S_{xx}\), and the intercept is \(a=\bar{y}-b\bar{x}\):

$$ b=\frac{47}{10}=4.7\text{ points per minute}, \qquad a=62-(4.7)(3)=47.9\text{ points}, \qquad \hat{y}=47.9+4.7x $$

The slope says that for each additional minute of practice, the fitted line predicts a quiz score that is 4.7 points higher, on average. This describes the linear association in these observations; it does not prove that an extra minute of practice causes a score increase. The intercept predicts a score of 47.9 points at zero minutes of practice. Since zero is outside the observed practice-time range of 1 to 5 minutes, treat that as an equation-based prediction rather than strong evidence about students who did no practice.

The sum of squared deviations in scores is \(S_{yy}=230\). Therefore, the correlation and coefficient of determination are:

$$ r=\frac{S_{xy}}{\sqrt{S_{xx}S_{yy}}} =\frac{47}{\sqrt{(10)(230)}}\approx 0.980, \qquad r^2\approx (0.980)^2\approx 0.960 $$

The positive \(r\), close to 1, summarizes a strong positive linear association between practice minutes and quiz score for these observations. The model’s \(r^2\) is about 0.960, or 96.0%. In this data set, about 96.0% of the variation in quiz scores is accounted for by the linear model relating score to practice time. This does not mean that 96.0% of each student’s score is explained, nor does it establish cause and effect.

To find \(s\), calculate fitted values and residuals. At \(x=1,2,3,4,5\), the fitted scores are \(52.6,57.3,62.0,66.7,71.4\). Subtracting these from the observed scores gives residuals \(-0.6,0.7,-1.0,2.3,-1.4\) points. The sum of squared residuals is \(0.36+0.49+1.00+5.29+1.96=9.10\). With \(n=5\):

$$ s=\sqrt{\frac{\sum (y-\hat{y})^2}{n-2}} =\sqrt{\frac{9.10}{3}} \approx 1.742\text{ points} $$

The observed scores are typically about 1.742 points vertically away from the fitted line. That is a response-scale description of residual scatter, not a guarantee that every prediction is within 1.742 points. Together, the slope describes the predicted rate of change, \(r\) summarizes direction and strength, \(r^2\) summarizes accounted-for variation, and \(s\) gives a typical residual size.

Worked Example: Explain a Meaningful Intercept

Worked Example: Audio Practice and Pronunciation Score

Original AP-style question. In an invented exercise, \(x\) is minutes of audio practice and \(y\) is the number of words pronounced correctly on a short task. The observed pairs are \((0,8),(1,10),(2,11),(3,13),(4,14)\). Interpret the fitted line’s slope and intercept, and state what the other numerical summaries contribute.

Solution. For these data, \(\bar{x}=2\), \(\bar{y}=11.2\), \(S_{xx}=10\), and \(S_{xy}=15\). The fitted line is:

$$ b=\frac{15}{10}=1.5\text{ words per minute}, \qquad a=11.2-(1.5)(2)=8.2\text{ words}, \qquad \hat{y}=8.2+1.5x $$

The slope means that the fitted score increases by 1.5 correctly pronounced words for each additional minute of practice, on average according to the line. The intercept predicts 8.2 correctly pronounced words at zero minutes. Zero is observed here, and a no-practice score makes sense in this setting, so the intercept has a direct interpretation as the fitted baseline. It is still a model prediction, not necessarily an observed score.

The sum of squared deviations in \(y\) is \(S_{yy}=22.8\), so \(r=15/\sqrt{(10)(22.8)}\approx 0.993\), and \(r^2\approx 0.987\). The correlation indicates a very strong positive linear association in these five observations. About 98.7% of the variation in pronunciation scores is accounted for by the fitted linear model.

The fitted scores are \(8.2,9.7,11.2,12.7,14.2\), giving residuals \(-0.2,0.3,-0.2,0.3,-0.2\) words. Their squared sum is \(0.30\), so \(s=\sqrt{0.30/(5-2)}=\sqrt{0.10}\approx 0.316\) words. This indicates that the observed scores are typically about 0.316 words from the fitted values. The small \(s\) and high \(r^2\) both describe close fit, but they express it differently: one in words and one as a proportion of variation.

The examples in this tutorial use small invented data sets to make the calculations visible. In an actual analysis, check the scatterplot and residual plot as well: numerical summaries alone can conceal curvature, changing spread, or unusual observations.

Worked Example: Choose Numbers for a Wait-Time Summary

Worked Example: Open Checkout Lanes and Customer Wait

Original AP-style question. In an invented store exercise, \(x\) is the number of checkout lanes open and \(y\) is a customer’s wait in minutes. Five observed pairs are \((1,48),(2,44),(3,43),(4,37),(5,38)\). A manager asks for numerical evidence describing the association, the typical residual size, and the model’s prediction at zero open lanes.

Solution. Here \(\bar{x}=3\), \(\bar{y}=42\), \(S_{xx}=10\), and \(S_{xy}=-27\). Thus:

$$ b=\frac{-27}{10}=-2.7\text{ minutes per lane}, \qquad a=42-(-2.7)(3)=50.1\text{ minutes}, \qquad \hat{y}=50.1-2.7x $$

The slope says that each additional open lane is associated with a predicted customer wait that is 2.7 minutes shorter, on average according to the fitted line. The correlation is \(r=-27/\sqrt{(10)(82)}\approx -0.942\), indicating a strong negative linear association. Squaring gives \(r^2\approx 0.889\), so about 88.9% of the variation in observed customer wait times is accounted for by this linear model.

The fitted waits for \(x=1,2,3,4,5\) are \(47.4,44.7,42.0,39.3,36.6\) minutes. The residuals are \(0.6,-0.7,1.0,-2.3,1.4\) minutes, with squared sum \(9.10\). Therefore \(s=\sqrt{9.10/(5-2)}\approx 1.742\) minutes. The residual standard deviation describes typical vertical scatter around the line in minutes.

The intercept predicts a wait of 50.1 minutes when zero lanes are open. But the observed range is 1 to 5 lanes, and a customer cannot enter a checkout lane when none are open. The intercept may be a feature of the fitted equation, but it is not a useful prediction for this situation. A clear report should include the negative slope and strong negative association, give \(s\) in minutes, and qualify the intercept rather than presenting it as a dependable real-world prediction.

Common Mistakes and What Full-Credit Communication Includes

  • Using \(r\) and \(r^2\) as if they say the same thing. \(r\) includes direction; \(r^2\) does not. State the sign and strength when reporting \(r\), and report \(r^2\) as the proportion of response variation accounted for by the linear model.
  • Calling \(r^2\) the percent of responses predicted correctly. It is about variation in the observed response values, not the percentage of individual responses the line gets right.
  • Giving a slope without units or context. “The slope is 4.7” is incomplete. Say that the predicted response changes by 4.7 points for each additional minute of practice, on average according to the fitted line.
  • Calling \(s\) a guaranteed prediction error. \(s\) is a typical residual size in response units. It does not say every prediction misses by exactly \(s\), or is within \(s\).
  • Interpreting every intercept as a meaningful real-world value. State that the intercept is the predicted response at \(x=0\), then consider whether zero is observed or makes sense in context. As in “Extrapolating Using the Intercept,” do not turn a mathematically defined intercept into an unsupported claim.
  • Claiming that a strong association proves cause and effect. Regression summaries describe an association in the data. Unless the study design supports a causal conclusion, use wording such as “is associated with” rather than “causes.”
  • Listing statistics without explaining what each adds. Connect every number to its purpose: direction and strength, explained variation, typical residual size, predicted change, or the fitted value at zero.

For a strong AP response, use the statistic that directly addresses the question, include its units when appropriate, and describe the population or data context carefully. Numerical evidence strengthens a model description when it agrees with the graphs and the setting; it does not replace those checks.

Key takeaway: Choose \(r\) for direction and strength, \(r^2\) for the proportion of response variation accounted for, \(s\) for typical residual size, the slope for predicted change per \(x\)-unit, and the intercept for the predicted response at \(x=0\). Interpret each in context and alongside the graphs.

Check Your Understanding

For each question, identify the most relevant regression number and describe what a complete interpretation should include.

  1. A report asks whether the linear association between two variables is positive or negative and how strong it is. Which statistic directly addresses this?
  2. A model has \(r^2=0.72\). What does this value say about variation in the response, and what does it not say?
  3. A fitted line predicts that the response changes by 3.4 units for every one-unit increase in \(x\). Which statistic is being interpreted, and what units should its value have?
  4. What does \(s=2.1\) tell you if the response is measured in minutes? Why is it not a guaranteed error bound?
  5. A model’s intercept is 15, but all observed \(x\)-values are between 20 and 60. What does the intercept mean mathematically, and what contextual caution is needed?
  6. Why is it useful to report both \(r^2\) and \(s\) when describing a model’s fit?