Tutorials › AP Statistics › Reading Full Regression Output Step by Step

Comparing and communicating regression models · Tutorial 988 of 1000

Reading Full Regression Output Step by Step

Follow a regression output table from its labels to a contextual interpretation of the equation, correlation, proportion of variation accounted for, and residual spread.

Intermediate 10 min read

What You'll Learn

  • Identify the intercept, slope, correlation, coefficient of determination, residual standard deviation, and sample size in a software report.
  • Distinguish the meanings and units of \(r^2\) and \(s\).
  • Translate a regression equation and its statistics into the context of the variables.
  • Use the observed \(x\)-range to judge whether a requested prediction is interpolation or extrapolation.
  • Avoid overinterpreting an intercept, a strong correlation, or a high \(r^2\).

What Does a Regression Output Table Tell You?

A fitted regression line is more than an equation on a screen. A software report may also display the correlation, \(r^2\), residual standard deviation, and number of cases. Reading those values together helps you describe the relationship, assess how closely the line fits, and decide what a prediction does—and does not—say.

As in “Linear Versus Nonlinear Model Choice,” the fitted line should be considered alongside the scatterplot, residual plot, and context. This tutorial focuses on reading the numerical output once a simple linear regression model has been fitted. Different programs may use different labels, but the core quantities have the same meanings.

Definition: In a simple linear regression report, the intercept \(a\) and slope \(b\) determine the fitted line \(\hat{y}=a+bx\). The correlation \(r\) describes the direction and strength of the linear association, \(r^2\) gives the proportion of variation in the response accounted for by the linear model, and \(s\) summarizes the typical size of the residuals in response units.

A software table does not decide what matters in a particular problem. You must identify which variable is explanatory, which is the response, their units, and the cases used to fit the line. Then interpret each statistic in that setting. In particular, \(r^2\) and \(s\) are not interchangeable: one is a proportion, while the other is measured in the response’s units.

Read the Report in a Consistent Order

Start by finding the equation and identifying what each coefficient represents. Some programs label the intercept \(a\) and slope \(b\); others label them “constant” and “coefficient.” Check the row labels rather than assuming the first number is the slope. The slope describes the predicted change in \(y\) for a one-unit increase in \(x\), with units of response per explanatory-variable unit. The intercept is the predicted response when \(x=0\).

Next, locate \(r\) and \(r^2\). The sign of \(r\) gives the direction of the linear association, and its magnitude summarizes its strength. In simple linear regression with an intercept, \(r^2\) is the square of \(r\). It describes the proportion of variation in the observed response values accounted for by the linear model. It does not give the percentage of cases predicted exactly, and it does not say that changing \(x\) causes a change in \(y\).

Find the report’s residual standard deviation, often labeled \(S\) by software. In this course, we write it as \(s\). It estimates the typical vertical distance between observed responses and the fitted line. Its units are the same as the response’s units. A value of \(s\) does not mean that every residual has that size; individual residuals can be smaller or larger.

Finally, note the number of cases and inspect the original data or output for the observed range of \(x\). The case count tells you how many observations were used, not whether they represent a broader population. The observed range lets you assess a proposed prediction: as discussed in “Reliability Within the Range of Data” and “Writing an Extrapolation Critique,” being beyond that range raises a separate concern from how well the line fits.

Reading checklist: Identify the variables and units; read the intercept and slope in the correct rows; interpret \(r\) and \(r^2\); interpret \(s\) in response units; record the number of cases; and check the observed \(x\)-range before using the equation for a prediction.

Worked Example: Reading a Complete Plant-Growth Report

Worked Example: Interpreting Seedling Height Output

Original AP-style question. A student records the number of days since transplanting, \(x\), and seedling height, \(y\), in centimeters, for five seedlings in an invented classroom activity. The observations are \((1,12),(2,15),(3,17),(4,20),(5,21)\). The software report is:

Output itemValue
Intercept \(a\)10.1000
Slope \(b\)2.3000
Correlation \(r\)0.9898
Coefficient of determination \(r^2\)0.9796
Residual standard deviation \(s\)0.6055 cm
Number of cases \(n\)5

State. Describe the fitted linear relationship between days since transplanting and seedling height, and explain the reported measures of association and fit.

Plan. Read the coefficients to write the fitted equation, then interpret \(r\), \(r^2\), and \(s\) using the variables and their units. The observations cover days 1 through 5, so any prediction should also be checked against that range.

Do. The fitted line is \(\hat{y}=10.1+2.3x\), where \(\hat{y}\) is predicted height in centimeters. For each additional day since transplanting, the model predicts an increase of 2.3 centimeters in height, on average across the fitted line. The intercept means that the model predicts a height of 10.1 centimeters at day 0. Day 0 is outside the observed range, so that intercept may be a mathematical part of the line without being a well-supported description of these seedlings at day 0.

The correlation, \(r=0.9898\), indicates a strong positive linear association between days since transplanting and height for these five seedlings. The value \(r^2=0.9796\) means that 97.96% of the variation in the observed seedling heights is accounted for by the fitted linear model. The residual standard deviation \(s=0.6055\) centimeters means that the observed heights typically differ from their fitted values by about 0.6055 centimeters. The report includes five cases.

The reported \(r^2\) is consistent with the reported correlation: \(0.9898^2\approx0.9797\), with the small difference due to rounding. The residuals from the equation are \(-0.4, 0.3, 0, 0.7,\) and \(-0.6\) centimeters. Their squared sum is \(0.16+0.09+0+0.49+0.36=1.10\). With \(n-2=3\) degrees of freedom for the residual standard deviation, \(s=\sqrt{1.10/3}\approx0.6055\) centimeters.

Conclude. For these five seedlings, the fitted line shows a strong positive linear association, and its residuals are typically about 0.6055 centimeters from the line. These descriptions apply to the observed seedlings; the output alone does not establish that time since transplanting caused the height differences or that the same relationship holds for other seedlings.

Worked Example: Use the Numbers, but Keep Their Meanings Separate

Worked Example: Interpreting Fan-Setting Output

Original AP-style question. An invented test records fan setting \(x\) and electrical power use \(y\), in watts, at settings 2, 4, 6, 8, and 10. The corresponding power-use measurements are 5, 9, 12, 16, and 18 watts. A software report gives \(a=2.10\), \(b=1.65\), \(r=0.9950\), \(r^2=0.9900\), \(s=0.6055\) watts, and \(n=5\). Interpret the output and predict power use at setting 7.

Solution. The fitted line is \(\hat{y}=2.10+1.65x\). For each one-unit increase in fan setting, predicted power use increases by 1.65 watts. The intercept predicts 2.10 watts at setting 0, but the observed settings range from 2 to 10. The output therefore does not provide observed data at setting 0 to support treating that intercept as an established power measurement there.

The correlation \(r=0.9950\) describes a very strong positive linear association in these five observations. The \(r^2\) value of 0.9900 means that 99.00% of the variation in the observed power-use measurements is accounted for by the fitted line. It does not mean that 99.00% of the predictions are exact. The residual standard deviation of 0.6055 watts describes the typical size of the observed-minus-predicted differences in power units.

For setting 7, substitute \(x=7\) into the fitted equation:

$$ \hat{y}=2.10+1.65(7)=13.65\text{ watts}. $$

Setting 7 lies between the observed settings 2 and 10, so this is an interpolation. The model predicts power use of 13.65 watts at that setting. This is a model prediction, not a guarantee that a particular fan will use exactly 13.65 watts. The strong \(r\) and high \(r^2\) summarize the linear association and fit; they do not eliminate individual prediction error or show that setting alone explains every difference in power use.

Worked Example: Match the Output to the Question

Worked Example: Predicting Resting Pulse From Training Weeks

Original AP-style question. In an invented set of observations, \(x\) is the number of weeks in a training program and \(y\) is resting pulse, in beats per minute. For five participants, the paired values are \((1,78),(2,74),(3,73),(4,69),(5,66)\). A software report gives \(a=80.7\), \(b=-2.9\), \(r=-0.9889\), \(r^2=0.9779\), and \(s=0.7958\) beats per minute. A reader claims, “Training reduces each person’s pulse by 2.9 beats per minute per week, and the model explains 97.79% of each person’s pulse.” Correct the interpretation and predict the fitted value at week 4.

Solution. The fitted equation is \(\hat{y}=80.7-2.9x\). The slope says that for each additional training week, the model predicts a decrease of 2.9 beats per minute in resting pulse, on average across the fitted line. It does not establish that training caused the decrease, nor does it say every participant’s pulse decreases by exactly that amount each week.

The negative correlation, \(r=-0.9889\), indicates a strong negative linear association between training weeks and resting pulse in these observations. The coefficient of determination, \(r^2=0.9779\), means that 97.79% of the variation in the observed resting-pulse values is accounted for by the fitted linear model. It does not mean the model explains 97.79% of each participant’s pulse. The value \(s=0.7958\) beats per minute describes the typical residual size in the response’s units.

At week 4, the fitted value is \(\hat{y}=80.7-2.9(4)=69.1\) beats per minute. Week 4 is within the observed range of 1 to 5 weeks. This is the line’s predicted response at that value of \(x\); it is not necessarily the observed pulse of any particular participant. The report summarizes an association in these five observations and, by itself, does not establish a cause-and-effect relationship.

Common Mistakes and AP Exam Tips

  • Reversing the intercept and slope. Read the row labels. A slope is the predicted response change for a one-unit increase in \(x\); an intercept is the predicted response at \(x=0\).
  • Leaving out units. State the slope in response units per explanatory-variable unit and \(s\) in response units. Correlation and \(r^2\) have no units.
  • Describing \(r^2\) as the percent of predictions that are correct. A full-credit interpretation says it is the proportion or percentage of variation in the observed response accounted for by the linear model, in context.
  • Confusing \(s\) with a typical percent error. \(s\) is measured in the response’s units and summarizes residual size; it is not a percentage and does not describe the error of every individual prediction exactly.
  • Giving the intercept an automatic practical meaning. The intercept refers to \(x=0\). Check whether zero is in the observed range and whether a response at zero is meaningful in context.
  • Treating strong association as causation or certainty. A large \(|r|\) or high \(r^2\) does not prove causation, guarantee accurate predictions for every case, or remove the need to check model fit and prediction range.

A complete interpretation ties each value to the variables. For example: “For each additional week of training, the fitted model predicts a decrease of 2.9 beats per minute in resting pulse. The model accounts for 97.79% of the variation in the observed pulse values, and residuals typically differ from the fitted values by about 0.7958 beats per minute.” That wording gives the slope’s direction and units, explains \(r^2\) correctly, and interprets \(s\) without promising exact predictions.

Key takeaway: Read the equation coefficients first, then interpret \(r\), \(r^2\), and \(s\) according to their different roles and units. Include the number of cases and check the observed \(x\)-range before predicting. Translate every relevant value into context without turning association into causation or a model summary into a guarantee.

Check Your Understanding

Use the meanings of the output values and their context to answer each question.

  1. A report gives \(a=6.2\) and \(b=0.8\), with \(x\) measured in hours and \(y\) measured in liters. State the slope’s meaning, including units.
  2. A model has \(r=-0.92\). What does the negative sign indicate, and what does the magnitude summarize?
  3. A report gives \(r^2=0.64\). Write a correct interpretation in context and identify one interpretation that would be incorrect.
  4. A software table labels the residual standard deviation \(S=3.1\), and the response is measured in points. What does \(S\) summarize, and what are its units?
  5. The observed explanatory-variable values run from 4 to 12. Is a prediction at \(x=13\) interpolation or extrapolation? What output detail alone cannot remove this concern?