From Summary Statistics to a Regression Equation
In “Linking the Slope to Correlation,” you learned that the least-squares regression slope can be calculated from the correlation and standard deviations. This tutorial uses that relationship to build a complete regression equation from summary statistics. First calculate the slope \(b\); then use \(b\) and the sample means to calculate the intercept \(a\).
The order matters. The slope calculation uses the correlation \(r\), the response standard deviation \(s_y\), and the predictor standard deviation \(s_x\). The intercept calculation then uses the slope and both sample means. Keep track of which variable is the predictor \(x\) and which is the response \(y\), because swapping them changes the calculations and the meaning of the line.
The first formula is the relationship from “Linking the Slope to Correlation.” The second formula rearranges the regression equation so that the intercept is calculated using the known mean response and the predicted response at the mean predictor value. Use the unrounded slope in this second calculation whenever possible, and round the final values only as appropriate for the context.
The slope’s units are response units per predictor unit, as explained in “Slope Units and Rates of Change.” The intercept has the same units as the response. The correlation has no units, so the standard-deviation ratio \(s_y/s_x\) supplies the slope’s units.
A Reliable Calculation Process
Before calculating, identify which summary statistics belong to which variable. A common source of error is placing the predictor standard deviation in the numerator. The slope must describe predicted change in \(y\) for a one-unit increase in \(x\), so use \(s_y\) over \(s_x\).
Multiply \(r\) by \(s_y/s_x\). Check that the sign agrees with \(r\), since the standard-deviation ratio is positive.
Substitute the slope and sample means into \(a=\bar{y}-b\bar{x}\). Keep the signs and units visible in the arithmetic.
Use \(\hat{y}=a+bx\). Confirm that the slope has response units per predictor unit and the intercept has response units.
The formulas require both variables to vary for \(r\) and the relationship \(b=r(s_y/s_x)\) to apply. If a variable is constant, do not use this calculation as though the correlation were defined. The special cases are described in “Linking the Slope to Correlation.”
Worked Calculations
Worked Example: Practice Time and Quiz Score
A fictional sample of students has \(x\) equal to minutes spent practicing a skill and \(y\) equal to a quiz score in points. The sample summaries are \(\bar{x}=6\) minutes, \(s_x=2\) minutes, \(\bar{y}=78\) points, \(s_y=8\) points, and \(r=0.75\). Calculate the least-squares regression equation.
First calculate the slope, placing the response standard deviation in the numerator:
The positive slope agrees with the positive correlation. Now use the slope and means to calculate the intercept:
The regression equation is \(\hat{y}=60+3x\), where \(x\) is practice time in minutes and \(\hat{y}\) is the predicted quiz score in points. In context, the line predicts a score 3 points higher for each additional minute of practice. The intercept says the line predicts 60 points at zero minutes of practice; whether that prediction is useful depends on whether zero minutes is meaningful and relevant to the data. This is a description of the fitted line, not a claim that practice alone causes a particular score.
As an arithmetic check on the intercept calculation, substitute the mean predictor value into the equation: \(60+3(6)=78\), the given mean response. This check can reveal a sign or multiplication error.
Worked Example: Outdoor Temperature and Heating Energy
A fictional building study records \(x\), the outdoor temperature in degrees Celsius, and \(y\), the building’s daily heating energy use in kilowatt-hours. Suppose \(\bar{x}=20\) degrees Celsius, \(s_x=5\) degrees Celsius, \(\bar{y}=42\) kilowatt-hours, \(s_y=12\) kilowatt-hours, and \(r=-0.50\). Find the regression equation and interpret its slope.
Calculate \(b\) from the correlation and standard deviations:
The negative sign matches the negative correlation. Next find the intercept, taking care that subtracting a negative product adds:
The equation is \(\hat{y}=66-1.2x\). In context, for each one-degree-Celsius increase in outdoor temperature, the line predicts 1.2 fewer kilowatt-hours of daily heating energy use. The intercept is the line’s predicted heating use at zero degrees Celsius. This example shows why writing the negative slope inside the intercept calculation is helpful: changing the sign incorrectly would produce a different intercept.
Check the means by evaluating the line at \(x=20\): \(66-1.2(20)=66-24=42\) kilowatt-hours. The result matches \(\bar{y}\). The slope’s units are kilowatt-hours per degree Celsius, while the intercept’s units are kilowatt-hours.
Worked Example: Reading Time and Pages Read
A fictional reading program records \(x\), a student’s average reading time per day in minutes, and \(y\), the number of pages the student reads in a week. The summaries are \(\bar{x}=30\) minutes, \(s_x=10\) minutes, \(\bar{y}=84\) pages, \(s_y=21\) pages, and \(r=0.40\). Calculate the slope and intercept, then write the equation.
Use \(s_y/s_x\) to convert the unitless correlation into the units of the slope:
Now substitute the slope and sample means into the intercept formula:
The regression equation is \(\hat{y}=58.8+0.84x\), with \(x\) measured in minutes per day and predicted \(y\) measured in pages per week. For each additional minute of average daily reading time, the line predicts 0.84 additional pages read per week. The intercept is the line’s predicted weekly page total for a student with zero average daily reading time. It is an extrapolation if zero is outside the range of reading times represented in the sample, so its practical interpretation should be made cautiously.
The equation also passes a simple numerical check at the sample mean of \(x\): \(58.8+0.84(30)=58.8+25.2=84\) pages. That agrees with the provided mean of \(y\).
Checks, Rounding, and Communication
A complete solution should show both stages of the calculation. Giving only a final equation makes it difficult to tell whether the correct standard deviations, means, and signs were used. Write the slope substitution first, then show the intercept substitution, and label the equation’s variables in context.
If the slope is a repeating decimal, retain extra digits while calculating \(a\). Rounding the slope too early can slightly change the intercept. For example, if a calculator gives a slope of \(1.2367\), use that value in \(a=\bar{y}-b\bar{x}\), and round the reported slope and intercept at the end. State that reported values are rounded when necessary.
The intercept calculation is sensitive to the sign of \(b\). If \(b\) is negative, then \(a=\bar{y}-b\bar{x}\) involves subtracting a negative number when \(\bar{x}\) is positive. Keeping parentheses around \(b\bar{x}\) is a useful safeguard. The means themselves can also be negative in some settings, so use the signed values as given rather than relying on a memorized pattern.
When interpreting the equation, follow the guidance in “Interpreting the Slope in Context” and “Interpreting the Y-Intercept in Context.” Describe the slope as the change in the line’s predicted response for a one-unit increase in the predictor. Describe the intercept as the line’s predicted response when the predictor is zero, then consider whether that value makes sense in context. Neither number guarantees what will happen for a particular individual.
Common Mistakes and AP Exam Tips
- Reversing the standard deviations. Using \(s_x/s_y\) instead of \(s_y/s_x\) gives the wrong scale and units. The numerator must be the standard deviation of the response.
- Using the wrong means in the intercept formula. The intercept is \(\bar{y}-b\bar{x}\), not \(\bar{x}-b\bar{y}\). Write the formula before substituting.
- Dropping a negative sign. If \(b\) is negative, keep it in parentheses in \(a=\bar{y}-b\bar{x}\). Subtracting a negative product may increase the intercept.
- Giving the correlation as the slope. The slope is \(r\) multiplied by \(s_y/s_x\). It usually differs from \(r\) and has units; \(r\) does not.
- Rounding before the final calculation. Use the calculator’s slope value, or several unrounded digits, to compute the intercept. Round the reported equation at the end.
- Reporting numbers without context. For full-credit communication, identify the predictor and response, provide the slope’s response-per-predictor units, and explain what the intercept represents at \(x=0\).
- Treating an intercept as automatically meaningful. The intercept is mathematically the predicted response at zero predictor value, but zero may be impossible, outside the data range, or otherwise unhelpful in context.
A strong calculation response makes the sequence clear: “Using \(b=r(s_y/s_x)\), the slope is [value] response units per predictor unit. Then \(a=\bar{y}-b\bar{x}=[value]\) response units, so the fitted equation is \(\hat{y}=a+bx\).” Add contextual interpretations when asked, and do not claim that the line describes a guaranteed change for every case.
Check Your Understanding
For each question, identify the predictor and response before substituting values. Both variables are assumed to vary unless stated otherwise.
- A sample has \(r=0.60\), \(s_x=4\), \(s_y=10\), \(\bar{x}=5\), and \(\bar{y}=30\). Calculate \(b\), then \(a\), and write the regression equation.
- A sample has \(r=-0.40\), \(s_x=5\), \(s_y=15\), \(\bar{x}=10\), and \(\bar{y}=50\). What are the slope and intercept?
- In the formula for \(b\), which standard deviation belongs in the numerator? Explain using the slope’s units.
- A student calculates \(b=-2\), \(\bar{x}=4\), and \(\bar{y}=10\). Calculate the intercept and show the signs in your substitution.
- Why should you use unrounded digits for the slope when calculating the intercept, if those digits are available?