Tutorials › AP Statistics › Regression Toward the Mean

Least-squares regression · Tutorial 875 of 1000

Regression Toward the Mean

Use the standardized regression equation to explain why an extreme parent height predicts a child height closer to the mean, while recognizing that individual children can differ.

Intermediate 9 min read

What You'll Learn

  • Explain regression toward the mean using the size of the correlation.
  • Use the standardized regression equation to compare predictor and predicted-response z-scores.
  • Calculate a parent-child height prediction in centimeters from sample summaries.
  • Distinguish a tendency in predicted values from a rule about every individual.
  • Describe how the strength and direction of a correlation affect predictions.
  • Identify why regression toward the mean is not, by itself, evidence of causation.

Why Extreme Values Lead to Less Extreme Predictions

In “Standardized Version of the Regression Line,” you learned that the predicted response z-score is \(\hat{z}_y=r z_x\). This equation gives a direct explanation for a familiar pattern: when the correlation is positive but less than 1, a predictor value far above its mean leads to a response prediction above its mean, but usually not as far above. A predictor value far below its mean leads to a prediction below its mean, but usually not as far below.

This pattern is called regression toward the mean. It refers to the regression line’s predictions: values of the predicted response tend to be less extreme, in standard-deviation units, than the corresponding predictor values when the variables have a positive correlation weaker than a perfect correlation. It does not mean the observed response for every individual must be closer to its mean.

Definition: Regression toward the mean is the tendency for a predicted response to be closer to its mean, in standard-deviation units, than an extreme predictor value is to its mean, when the variables have a positive correlation less than 1. In standardized form, \(\hat{z}_y=r z_x\), so if \(0<r<1\), then \(|\hat{z}_y|<|z_x|\) for a nonzero \(z_x\).

The key is the size of \(r\). A positive correlation means the prediction is on the same side of its mean as the predictor. Because a correlation has absolute value no greater than 1, multiplying by a positive \(r\) smaller than 1 reduces the predictor’s distance from zero on the standardized scale. If \(r=1\), the predicted z-score is just as extreme as the predictor z-score. If \(r=0\), the predicted response z-score is 0 for every predictor value: the line predicts the response mean.

The word “regression” here does not mean that a person, family, or population necessarily changes or improves. It describes a feature of predictions from a regression line. The line captures the overall linear pattern, while individual responses vary around that pattern.

A Parent-Child Height Example

Consider an invented example involving adult parents and their adult children of the same gender. Let \(x\) be a parent’s height and \(y\) be the adult child’s height, both measured in centimeters. Suppose a fictional sample has \(\bar{x}=170\) cm, \(s_x=8\) cm, \(\bar{y}=170\) cm, \(s_y=7\) cm, and \(r=0.60\). These summaries are chosen for illustration; they are not results from a real study.

A parent who is much taller than the sample mean has a positive \(z_x\). Since \(r=0.60\), the predicted child z-score is positive too, but only 60% as large as the parent’s z-score. The same pattern applies in the other direction: a very short parent leads to a below-average predicted child height, but the prediction is not as far below the child-height mean in standard-deviation units.

Worked Example: Predicting a Child’s Height From a Tall Parent

Using the fictional parent-child height summaries above, predict the adult child’s height for a parent who is 186 cm tall. Explain how the prediction illustrates regression toward the mean.

State. Parent height is the predictor \(x\), and adult child height is the response \(y\). We want the line’s predicted child height, not a guarantee of the actual child’s height.

Plan. Find the parent’s z-score using the parent mean and standard deviation. Then use \(\hat{z}_y=r z_x\) and convert the predicted child z-score to centimeters using the child mean and standard deviation. As a check, calculate the original regression line from the summaries.

Do: standardized calculation. The parent is 16 cm above the parent-height mean, or two parent standard deviations above it:

$$ z_x=\frac{186-170}{8}=2. $$

The predicted child z-score is:

$$ \hat{z}_y=(0.60)(2)=1.20. $$

Convert this predicted z-score to centimeters:

$$ \hat{y}=170+(1.20)(7)=170+8.4=178.4\text{ cm}. $$

Check: original regression line. Using \(b=r(s_y/s_x)\), the slope is \(0.60(7/8)=0.525\) centimeters of predicted child height per centimeter of parent height. The intercept is \(a=\bar{y}-b\bar{x}=170-(0.525)(170)=80.75\) centimeters. Thus, \(\hat{y}=80.75+0.525x\), and at \(x=186\):

$$ \hat{y}=80.75+(0.525)(186)=80.75+97.65=178.4\text{ cm}. $$

Conclude. For a parent 186 cm tall, the regression line predicts an adult child height of 178.4 cm. The parent is 2 standard deviations above the parent-height mean, while the predicted child is 1.2 standard deviations above the child-height mean. The predicted child is still taller than average, but the prediction is less extreme relative to its own mean.

What Happens at the Low End?

Regression toward the mean works symmetrically in this positive-association example. A parent who is unusually short relative to the parent-height sample mean has a negative z-score. Multiplying by a positive correlation gives a negative predicted child z-score as well, but with a smaller absolute value when \(0<r<1\). The prediction remains below the child-height mean, just not as far below in standard-deviation units.

Worked Example: Predicting a Child’s Height From a Short Parent

Use the same fictional summaries: \(\bar{x}=\bar{y}=170\) cm, \(s_x=8\) cm, \(s_y=7\) cm, and \(r=0.60\). Predict the adult child’s height for a parent who is 154 cm tall, and compare the standardized distances from the means.

State. The parent is the predictor and the adult child is the response. The parent’s height is below its sample mean, so we expect a below-mean prediction because the correlation is positive.

Plan. Calculate \(z_x\), multiply by \(r\), and convert the predicted response z-score to centimeters. Compare the absolute values of the two z-scores to assess whether the prediction is less extreme.

Do. The parent is 16 cm below the parent-height mean:

$$ z_x=\frac{154-170}{8}=-2. $$

The predicted child z-score is:

$$ \hat{z}_y=(0.60)(-2)=-1.20. $$

In centimeters, the prediction is:

$$ \hat{y}=170+(-1.20)(7)=170-8.4=161.6\text{ cm}. $$

The parent is 2 standard deviations below the parent mean, while the predicted child is 1.2 standard deviations below the child mean. Equivalently, the absolute standardized distances are \(2\) and \(1.2\), so the predicted response is closer to its mean.

Conclude. For a parent who is 154 cm tall, the line predicts an adult child height of 161.6 cm. This prediction is below the sample mean child height, but it is less extreme in standard-deviation units than the parent’s height is relative to the parent mean.

A Prediction Is Not a Guarantee

The word “tendency” matters. A regression line predicts an average response for a given predictor value; it does not say that every observed response will equal its prediction. In the parent-child example, a tall parent can have an adult child who is taller than the line predicts. Another tall parent can have a child who is shorter than predicted. As covered in “Prediction Versus Observed Values,” the difference between an observed response and its predicted value is the residual, \(y-\hat{y}\).

For example, a parent 182 cm tall is \(1.5\) standard deviations above the parent-height mean. The line predicts a child height \(0.60(1.5)=0.90\) standard deviations above the child-height mean, or \(170+(0.90)(7)=176.3\) cm. If one fictional child in the sample is actually 181 cm tall, the residual is \(181-176.3=4.7\) cm. That child’s actual height is more extreme than the line’s prediction, but this does not contradict regression toward the mean. The concept describes the prediction, not a rule that every observed child must be less extreme than their parent.

The same distinction applies if a person is selected because of an extreme response. For instance, a student chosen for an unusually high score on one test might not score as far above average on another test. A regression line can predict a less extreme score when the two measurements are positively associated but imperfectly correlated. That is a prediction pattern, not proof that the first score caused the second score to change.

Worked Example: A Child’s Actual Height Can Differ From the Prediction

For the fictional parent-child data, suppose a 182 cm parent has an adult child who is actually 181 cm tall. The sample summaries are \(\bar{x}=\bar{y}=170\) cm, \(s_x=8\) cm, \(s_y=7\) cm, and \(r=0.60\). Find the predicted height and residual, then explain what the comparison shows.

State. The predictor is parent height, the response is the child’s observed height, and the regression line prediction will be compared with that observed response.

Plan. Standardize the parent height, use \(\hat{z}_y=r z_x\), and convert the prediction to centimeters. Then calculate the residual as observed child height minus predicted child height.

Do. The parent’s z-score is:

$$ z_x=\frac{182-170}{8}=1.5. $$

The predicted child z-score and height are:

$$ \hat{z}_y=(0.60)(1.5)=0.90, \qquad \hat{y}=170+(0.90)(7)=176.3\text{ cm}. $$

The residual is:

$$ y-\hat{y}=181-176.3=4.7\text{ cm}. $$

Conclude. The line predicts a child height of 176.3 cm, while the child’s observed height is 181 cm, giving a positive residual of 4.7 cm. The actual child is taller than predicted. The prediction is still less extreme than the parent’s height in standardized units, but this individual outcome is not itself required to be less extreme.

How Correlation Strength Affects the Pattern

The closer a positive \(r\) is to 1, the less the prediction is pulled toward the response mean in standardized units. For example, if \(z_x=2\), then \(r=0.90\) gives \(\hat{z}_y=1.80\), while \(r=0.40\) gives \(\hat{z}_y=0.80\). Both predictions are above the response mean, but the first is more extreme. If \(r=0\), the line predicts the response mean, regardless of how far the predictor is from its mean.

A negative correlation changes the direction of the prediction: a predictor above its mean leads to a predicted response below its mean. When \(-1<r<0\), multiplying by \(r\) also reduces the absolute standardized distance. This course’s usual parent-child height illustration has a positive association, so the predictor and predicted response are on the same side of their respective means.

Remember that standardized distances are relative to each variable’s own mean and standard deviation. Saying a parent is 2 standard deviations above the parent mean does not mean a child prediction is 2 centimeters, or even the same number of centimeters, from the child mean. Use \(s_x\) to standardize parent height and \(s_y\) to convert the predicted child z-score back to centimeters.

Common Mistakes and AP Exam Tips

  • Claiming every child must be closer to the mean than the parent. Regression toward the mean describes the line’s prediction, not every observed parent-child pair. A full-credit response distinguishes the predicted child height from an individual child’s actual height.
  • Comparing raw centimeter differences when the standard deviations differ. The claim that the prediction is less extreme is made in standard-deviation units. Compare \(|z_x|\) with \(|\hat{z}_y|\), not simply the raw distances in centimeters.
  • Forgetting the sign of \(r\). Use \(r\), not \(|r|\), in \(\hat{z}_y=r z_x\). A negative correlation reverses the direction of the predicted response relative to the response mean.
  • Calling regression toward the mean a causal effect. A regression line summarizes an association and makes predictions. It does not show that a parent’s height causes a child’s height to move toward an average.
  • Using the wrong variable’s standard deviation. Standardize parent height with \(s_x\); convert a predicted child z-score with \(s_y\). Each variable has its own mean and standard deviation.
  • Overstating a prediction outside the observed range. As discussed in “Predicting Within the Data Range,” a prediction for an extreme predictor value beyond the observed range is extrapolation and may be unreliable. Check the data range before treating the prediction as useful.

For an AP-style explanation, name the predictor and response, state what the line predicts, and compare the distances in standard-deviation units. For example: “Because the correlation between parent and adult child heights is positive but less than 1, a parent who is 2 standard deviations above the parent mean has a predicted child height only 1.2 standard deviations above the child mean. This is a tendency in the regression prediction, not a guarantee about an individual child.”

Key takeaway: When \(0<r<1\), the equation \(\hat{z}_y=r z_x\) makes the predicted response less extreme than the predictor in standardized units. Regression toward the mean describes the line’s predictions, not a rule for every observed case or evidence of causation.

Check Your Understanding

Use the standardized regression equation where needed. Interpret predictions as predictions from a line, not guarantees about individuals.

  1. A positive correlation is \(r=0.70\), and a case has \(z_x=2.5\). Find \(\hat{z}_y\). Is the predicted response more or less extreme than the predictor in standard-deviation units?
  2. For a fictional parent-child height model, \(r=0.50\), \(\bar{y}=168\) cm, and \(s_y=6\) cm. A parent’s height is 2 standard deviations above the parent mean. Find the predicted child z-score and predicted height.
  3. Explain why an actual child can be taller than the predicted height without disproving regression toward the mean.
  4. If \(r=1\) and \(z_x=-1.8\), what is \(\hat{z}_y\)? Does this example show regression toward the mean?
  5. Why should an AP response compare standardized distances rather than conclude that the prediction is closer to the mean based only on raw centimeter distances?