Tutorials › AP Statistics › Why Regressing Y on X Differs From X on Y

Least-squares regression · Tutorial 870 of 1000

Why Regressing Y on X Differs From X on Y

Understand why reversing the predictor and response creates a different least-squares regression line, and practice choosing the line that answers the question.

Intermediate 10 min read

What You'll Learn

  • Distinguish a regression of y on x from a regression of x on y.
  • Calculate both regression slopes using the correlation and standard deviations.
  • Explain why the two slopes are not generally reciprocals.
  • Identify why the two regression lines pass through the sample means.
  • Choose the response variable to match the prediction question.
  • Compare a separate reverse regression with the algebraic inverse of a line.

Why the Predictor and Response Roles Matter

In “Predicting x From y Using the Equation,” you saw that solving a regression equation for the predictor is not generally the same as fitting a new regression with the roles reversed. The reason is built into the least-squares criterion: a regression line is fitted to predict one variable from another, and it minimizes squared residuals in the response variable.

A regression of \(y\) on \(x\) predicts \(y\) from \(x\), so its equation is \(\hat{y}=a+bx\). A regression of \(x\) on \(y\) predicts \(x\) from \(y\), so its equation is \(\hat{x}=c+dy\). These are different fitting tasks. The predictor is not merely a label: it determines which variable’s values are used to predict the other and which residuals are squared.

Definition: The regression of \(y\) on \(x\) treats \(x\) as the predictor and \(y\) as the response. The regression of \(x\) on \(y\) treats \(y\) as the predictor and \(x\) as the response. Each least-squares line minimizes the sum of squared residuals in its response variable.

As in “The Idea of the Least-Squares Criterion,” the residual for a case is its observed response minus its predicted response. Thus, the \(y\)-on-\(x\) line minimizes squared differences between observed \(y\)-values and predicted \(\hat{y}\)-values. The \(x\)-on-\(y\) line minimizes squared differences between observed \(x\)-values and predicted \(\hat{x}\)-values. On a scatterplot with \(x\) on the horizontal axis and \(y\) on the vertical axis, these correspond to vertical and horizontal prediction errors, respectively.

The Two Slopes

When both variables vary, the slope formula from “Linking the Slope to Correlation” gives the slope for predicting \(y\) from \(x\). Reversing the roles means using \(y\) as the predictor instead. In each case, the slope has response units per predictor unit.

Formula: For paired data with correlation \(r\), sample standard deviations \(s_x\) and \(s_y\), and both variables varying:
$$ \begin{aligned} b_{y\text{ on }x}&=r\frac{s_y}{s_x},\\ d_{x\text{ on }y}&=r\frac{s_x}{s_y}. \end{aligned} $$
The first slope has \(y\)-units per \(x\)-unit. The second has \(x\)-units per \(y\)-unit.

The formulas use the same \(r\), but the standard-deviation ratios are reversed. The slopes are generally not reciprocals. Their product is:

$$ b_{y\text{ on }x}d_{x\text{ on }y} =\left(r\frac{s_y}{s_x}\right)\left(r\frac{s_x}{s_y}\right)=r^2. $$

By contrast, algebraically solving \(\hat{y}=a+bx\) for \(x\) gives a line with slope \(1/b\), provided \(b\ne0\). Since the reverse-regression slope is \(d=r s_x/s_y\), while \(1/b=s_x/(r s_y)\), these are generally different. They match only when \(r^2=1\), the case of a perfect linear association. If \(r=0\), the \(y\)-on-\(x\) line is horizontal and cannot be inverted to obtain a unique \(x\).

Both regression lines pass through the point \((\bar{x},\bar{y})\), as covered in “Finding the Line Through the Means.” That shared point does not make the lines identical. Away from the means, they generally give different predictions because they were fitted to minimize different residuals.

Choose the Line That Matches the Question

The wording of a prediction question identifies the response. If the question gives an \(x\)-value and asks for a predicted \(y\), use the \(y\)-on-\(x\) line. If it gives a \(y\)-value and asks for a predicted \(x\), use the \(x\)-on-\(y\) line. This is a choice about what the model is meant to predict, not a choice based on which equation looks easier to rearrange.

1
Identify what is being predicted.
The variable the question asks you to estimate is the response.
2
Assign the predictor and response.
Use the given variable as the predictor and the requested variable as the response.
3
Use the matching regression line.
Use \(\hat{y}=a+bx\) to predict \(y\) from \(x\), or \(\hat{x}=c+dy\) to predict \(x\) from \(y\).
4
Check the interpretation and units.
Report the predicted response in its own units. Do not call the algebraic inverse a separately fitted regression.

Worked Examples

Worked Example: Calculate Both Slopes From Summary Statistics

A fictional plant study records plant height \(x\), in centimeters, and leaf mass \(y\), in grams. The sample summaries are \(r=0.60\), \(s_x=4\) centimeters, \(s_y=10\) grams, \(\bar{x}=20\) centimeters, and \(\bar{y}=35\) grams. Find both regression equations.

For the regression predicting leaf mass from height, the response is \(y\). Its slope is:

$$ b_{y\text{ on }x} =0.60\left(\frac{10\text{ g}}{4\text{ cm}}\right) =1.5\text{ g per cm}. $$

The intercept is \(\bar{y}-b\bar{x}=35-1.5(20)=5\) grams. Therefore, the line is \(\hat{y}=5+1.5x\), with the slope measured in grams per centimeter.

For the regression predicting height from leaf mass, the response is now \(x\). Its slope and intercept are:

$$ d_{x\text{ on }y} =0.60\left(\frac{4\text{ cm}}{10\text{ g}}\right) =0.24\text{ cm per g}, \qquad c=\bar{x}-d\bar{y}=20-0.24(35)=11.6\text{ cm}. $$

The separate reverse-regression equation is \(\hat{x}=11.6+0.24y\). The slopes are not reciprocals: their product is \(1.5(0.24)=0.36\), which equals \(r^2=0.60^2=0.36\). The reciprocal of 1.5 is about 0.6667, not 0.24.

Both equations agree at the means. The first predicts \(5+1.5(20)=35\) grams at a height of 20 centimeters; the second predicts \(11.6+0.24(35)=20\) centimeters at a leaf mass of 35 grams. The units also confirm which variable each line predicts.

Worked Example: Compare a Reverse Regression With an Algebraic Inverse

A fictional class records weekly study hours \(x\) and quiz scores \(y\), in points:

Study hours, \(x\)Quiz score, \(y\)
152
258
355
467
568

The means are \(\bar{x}=3\) hours and \(\bar{y}=60\) points. The deviations from the means are \(-2,-1,0,1,2\) for \(x\), and \(-8,-2,-5,7,8\) for \(y\). Their cross-products sum to:

$$ \sum (x-\bar{x})(y-\bar{y}) =(-2)(-8)+(-1)(-2)+(0)(-5)+(1)(7)+(2)(8) =16+2+0+7+16=41. $$

The sum of squared \(x\)-deviations is \(4+1+0+1+4=10\), and the sum of squared \(y\)-deviations is \(64+4+25+49+64=206\). Therefore, the \(y\)-on-\(x\) slope is \(41/10=4.1\) points per hour. Its intercept is \(60-4.1(3)=47.7\) points, giving \(\hat{y}=47.7+4.1x\).

For the \(x\)-on-\(y\) regression, the predictor is quiz score, so the slope is \(41/206\approx0.1990\) hours per point. Its intercept is:

$$ 3-\left(\frac{41}{206}\right)(60) =\frac{618-2460}{206} =-\frac{921}{103} \approx -8.9417\text{ hours}. $$

Thus, the reverse line is \(\hat{x}\approx -8.9417+0.1990y\). As a check on the slopes, their product is \(4.1(41/206)=168.1/206\approx0.8160\). The correlation calculated from these sums is \(r=41/\sqrt{10(206)}\approx0.9033\), so \(r^2\approx0.8160\), as the slope-product formula predicts.

Suppose the question asks for the study time associated with a predicted quiz score of 64 points. Solving the first line algebraically gives:

$$ x=\frac{64-47.7}{4.1} =\frac{16.3}{4.1} \approx3.9756\text{ hours}. $$

The separate \(x\)-on-\(y\) regression instead predicts:

$$ \hat{x}=-\frac{1842}{206}+\frac{41}{206}(64) =\frac{-1842+2624}{206} =\frac{391}{103} \approx3.7961\text{ hours}. $$

These predictions differ. The first is the algebraic inverse of the line fitted to predict quiz score; the second comes from a line fitted to predict study time. If the question is specifically asking for predicted study hours given a quiz score, the second regression is the matching model.

Worked Example: Match a Negative Association to the Prediction Goal

A fictional building dataset relates outdoor temperature \(x\), in degrees Celsius, to daily heating energy \(y\), in kilowatt-hours. Suppose \(r=-0.75\), \(s_x=8\) degrees Celsius, \(s_y=24\) kilowatt-hours, \(\bar{x}=12\) degrees Celsius, and \(\bar{y}=80\) kilowatt-hours.

If the goal is to predict energy use from temperature, calculate the slope and intercept for \(y\) on \(x\):

$$ b_{y\text{ on }x} =-0.75\left(\frac{24}{8}\right) =-2.25\text{ kilowatt-hours per degree Celsius}, \qquad a=80-(-2.25)(12)=107. $$

The equation is \(\hat{y}=107-2.25x\). For example, at \(x=16\) degrees Celsius, it predicts \(107-2.25(16)=107-36=71\) kilowatt-hours.

If the goal is instead to predict temperature from energy use, the slope is \(-0.75(8/24)=-0.25\) degrees Celsius per kilowatt-hour. The intercept is \(12-(-0.25)(80)=32\) degrees Celsius, so \(\hat{x}=32-0.25y\). At \(y=71\) kilowatt-hours, this line predicts \(32-0.25(71)=14.25\) degrees Celsius, not 16 degrees.

The difference is not a sign error: both slopes are negative, as they should be for this negative correlation, but the lines answer different questions. The appropriate one depends on which measurement is the response being predicted.

Common Mistakes and AP Exam Tips

  • Reversing the labels but not the formula. For \(y\) on \(x\), use \(r(s_y/s_x)\). For \(x\) on \(y\), use \(r(s_x/s_y)\). Always identify the response before choosing the formula.
  • Assuming the slopes are reciprocals. Their product is \(r^2\), not generally 1. The reciprocal slope comes from algebraically rearranging one line, not from fitting the other regression.
  • Giving the slope the wrong units. A slope has response units per predictor unit. State grams per centimeter for leaf mass predicted from height, for example—not centimeters per gram.
  • Thinking the shared mean point means the lines are identical. Both lines pass through \((\bar{x},\bar{y})\), but can have different slopes and predictions elsewhere.
  • Choosing a regression based on which variable seems more important. Use the roles specified by the question: the variable being predicted is the response. Do not make a causal claim from the regression alone.
  • Calling a rearranged equation a reverse regression. A reverse regression is fitted separately with the variables’ roles exchanged. Make that distinction explicit in your explanation.

For a strong AP-style response, name the predictor and response, use the matching equation, include the slope’s units, and explain what the prediction means in context. If comparing equations, state that the reverse regression minimizes squared residuals in \(x\), while the original regression minimizes squared residuals in \(y\).

Key takeaway: Regressing \(y\) on \(x\) and regressing \(x\) on \(y\) are different least-squares tasks. Their slopes are \(r(s_y/s_x)\) and \(r(s_x/s_y)\), so the lines are generally different. Choose the response variable to match what the question asks you to predict.

Check Your Understanding

For each question, keep the predictor, response, and units clear.

  1. If \(r=0.50\), \(s_x=3\) hours, and \(s_y=6\) hours of sleep, find both regression slopes and include the units of each.
  2. What does the \(y\)-on-\(x\) regression minimize? What does the \(x\)-on-\(y\) regression minimize?
  3. If the slope predicting \(y\) from \(x\) is 2.4 and \(r=0.80\), what is the slope for predicting \(x\) from \(y\)? Explain using the product of the slopes.
  4. Both regression lines pass through the sample means. Why does that not mean they make identical predictions?
  5. A question gives a quiz score and asks for predicted study time. Which variable is the response, and which regression should be used?