Predicting Relative Standing With a Regression Line
In “Properties of the Least-Squares Line,” you learned that the line passes through \((\bar{x},\bar{y})\), and in “Linking the Slope to Correlation” you saw that its slope is \(b=r(s_y/s_x)\). These facts lead to a useful way to express a regression prediction: instead of predicting the response in its original units, predict how many standard deviations above or below its mean the response is predicted to be.
To do this, standardize both variables. A z-score describes a value’s distance from its variable’s mean in standard-deviation units. The standardized regression equation is especially simple: the predicted response z-score is the correlation multiplied by the predictor z-score. This lets you predict relative standing directly, without first calculating the line’s intercept and slope in the original units.
Why the Standardized Equation Works
The equation follows from two properties of the least-squares line already established in this course. First, the line passes through \((\bar{x},\bar{y})\). Second, its slope is \(b=r(s_y/s_x)\). Therefore, the predicted response can be written as the mean response plus the slope times the predictor’s distance from its mean:
To express that prediction in response standard-deviation units, subtract \(\bar{y}\) and divide by \(s_y\). Then use \(b=r(s_y/s_x)\) and \(z_x=(x-\bar{x})/s_x\):
This is the standardized version of the regression line. Its intercept is \(0\), because a predictor at its mean has \(z_x=0\) and the line predicts a response at its mean, \(\hat{z}_y=0\). Its slope is \(r\). The correlation is unitless, and standardized values are measured in standard-deviation units, so the equation does not involve the original measurement units.
For example, if \(z_x=1\) and \(r=0.70\), the line predicts \(\hat{z}_y=0.70\): the response is predicted to be \(0.70\) standard deviations above its mean. This is a statement about relative standing, not about the response’s original units. To obtain a prediction in those units, multiply \(\hat{z}_y\) by \(s_y\) and add \(\bar{y}\).
Worked Example: Predicting a Response’s Relative Standing
In a fictional study of students, \(x\) is weekly time spent on a reading app, and \(y\) is a reading-comprehension score. The correlation is \(r=0.80\). One student’s weekly app time is \(1.5\) standard deviations above the sample mean, so \(z_x=1.5\). Find and interpret the response’s predicted z-score.
State. The predictor is app time, and the response is the comprehension score. We want the score’s predicted relative standing, not yet a prediction in score points.
Plan. Use the standardized regression equation \(\hat{z}_y=r z_x\). Then interpret the sign and size of the predicted z-score in context.
Do.
Conclude. The line predicts a comprehension score \(1.20\) standard deviations above the sample mean score for a student whose weekly app time is \(1.5\) standard deviations above the sample mean app time. This is a predicted relative standing; it does not say the student’s actual score must be exactly \(1.20\) standard deviations above the mean.
Positive and Negative Relative Standing
The sign of \(r\) determines how the line relates the predictor’s relative standing to the response’s predicted relative standing. When \(r\) is positive, predictor values above their mean have positive z-scores and produce positive predicted response z-scores. Predictor values below their mean produce negative predicted response z-scores. When \(r\) is negative, the signs go in opposite directions: a predictor value above its mean leads to a prediction below the response mean, and vice versa.
The size of the correlation matters too. Since correlations have absolute value at most \(1\), multiplying by \(r\) cannot make the predicted response z-score larger in absolute value than the predictor z-score. For instance, if \(r=-0.60\) and \(z_x=1.5\), then \(\hat{z}_y=-0.90\). The line predicts a response below its mean, and the predicted distance is \(0.90\) response standard deviations. The multiplication accounts for both direction and size.
This rule concerns the predicted response. An actual response can lie above or below its prediction; as covered in “Prediction Versus Observed Values,” the difference between an observed response and its prediction is the residual. Thus, a predicted z-score is not a guarantee about the actual z-score for a particular case.
Worked Example: A Negative Correlation and a Predictor Below Its Mean
A fictional technology comparison examines a device’s screen-brightness setting \(x\) and its battery duration \(y\). The sample correlation is \(r=-0.64\). A device’s brightness setting is \(1.25\) standard deviations below the sample mean, so \(z_x=-1.25\). Find and interpret the predicted battery-duration z-score.
State. The predictor is screen brightness and the response is battery duration. Because the correlation is negative, the line predicts response standing in the opposite direction from predictor standing.
Plan. Multiply \(r\) by \(z_x\). A positive product will mean the predicted battery duration is above its sample mean.
Do.
Conclude. The line predicts a battery duration \(0.80\) standard deviations above the sample mean for a device with a brightness setting \(1.25\) standard deviations below the sample mean. Both z-scores are relative to their own variable’s mean and standard deviation.
Converting a Predicted Z-Score Back to Original Units
The standardized equation is useful when a question asks for relative standing. If the question instead asks for a predicted response in its original units, convert the predicted z-score back. The definition \(\hat{z}_y=(\hat{y}-\bar{y})/s_y\) rearranges to:
You can also find the predictor z-score from its original value using \(z_x=(x-\bar{x})/s_x\). This gives a complete path from an original predictor value to a predicted response: standardize \(x\), multiply by \(r\), and convert the predicted response z-score back using \(\bar{y}\) and \(s_y\). Keep track of which mean and standard deviation belong to which variable.
Worked Example: Checking a Standardized Prediction Against the Original Line
A fictional plant-growth project records daily sunlight \(x\), in hours, and plant height \(y\), in centimeters. The sample summaries are \(\bar{x}=5\) hours, \(s_x=1.5\) hours, \(\bar{y}=24\) centimeters, \(s_y=6\) centimeters, and \(r=0.60\). Find the predicted plant height when daily sunlight is \(6.5\) hours. Show the standardized calculation and check it using the original regression line.
State. Sunlight is the predictor and plant height is the response. The requested prediction is a height in centimeters, so we will convert the predicted z-score back to centimeters.
Plan. First find \(z_x\), then use \(\hat{z}_y=r z_x\). Convert using \(\hat{y}=\bar{y}+\hat{z}_y s_y\). As a check, calculate the slope and intercept from the summaries and predict with \(\hat{y}=a+bx\).
Do: standardized calculation. The predictor is \(1.5\) hours above its mean, which is one predictor standard deviation:
The predicted response z-score is:
Convert this to centimeters using the response mean and standard deviation:
Check: original regression line. The slope is \(b=r(s_y/s_x)\), and the intercept is \(a=\bar{y}-b\bar{x}\):
Thus, the original line is \(\hat{y}=12+2.4x\). At \(x=6.5\) hours, it gives:
Conclude. Both methods predict a plant height of \(27.6\) centimeters. In relative-standing terms, that prediction is \(0.60\) standard deviations above the sample mean plant height. The agreement is a useful check that the standardization and conversion used the correct summaries.
Common Mistakes and AP Exam Tips
- Writing \(z_y=r z_x\) as though it describes every observed pair. The regression equation is \(\hat{z}_y=r z_x\). The hat matters: it marks the response value predicted by the line. An observed response can differ from that prediction.
- Using the response standard deviation to standardize \(x\). Calculate \(z_x\) with \(\bar{x}\) and \(s_x\), and convert a predicted response using \(\bar{y}\) and \(s_y\). Each variable has its own mean and standard deviation.
- Dropping the sign. Do not use \(|r|\) in the equation. A negative correlation reverses the direction of predicted relative standing. Keep both signs when multiplying.
- Calling a z-score a percentage or a percentile. A z-score is a number of standard deviations from the mean; it is not itself a percent or a percentile rank. State the predicted distance above or below the mean in standard-deviation units.
- Confusing the correlation with a predicted response z-score. The predicted response z-score equals \(r\) only when \(z_x=1\). In general, calculate \(r z_x\).
- Stopping at the z-score when the question asks for original units. Convert with \(\hat{y}=\bar{y}+\hat{z}_y s_y\). Include the response’s units in the final prediction.
- Overstating what the line predicts. A prediction is not a promise that an individual’s actual response will equal \(\hat{y}\). A complete interpretation says what the line predicts, in context, and does not claim that every case lies on the line.
For an AP-style response, state whether the predicted response is above or below its mean, give the predicted distance in standard-deviation units, and identify the relevant context. For example: “For a student whose weekly app time is \(1.5\) standard deviations above the sample mean, the regression line predicts a comprehension score \(1.20\) standard deviations above the sample mean score.” If original units are requested, also report the converted prediction and its units.
Check Your Understanding
Use the standardized regression equation to answer each question. Treat each predicted z-score as a prediction from the regression line.
- A data set has \(r=0.75\). For a case with \(z_x=2\), what is \(\hat{z}_y\), and how should it be described relative to the response mean?
- A data set has \(r=-0.50\). For a case with \(z_x=-1.4\), calculate and interpret \(\hat{z}_y\).
- In a fictional study, \(r=0.40\), \(\bar{y}=80\) points, and \(s_y=10\) points. For a case with \(z_x=1.5\), find the predicted response z-score and the predicted response in points.
- Explain why the equation uses \(\hat{z}_y\), rather than claiming that every observed response has z-score \(r z_x\).
- A student calculates \(z_x\) using the response standard deviation. Identify the error and state which standard deviation should be used.