Tutorials › AP Statistics › Transforming Data to Achieve Linearity

Unusual points and model fit · Tutorial 977 of 1000

Transforming Data to Achieve Linearity

See how to recognize when a logarithmic transformation improves a linear model and how to use the transformed model to make predictions in the original units.

Intermediate 11 min read

What You'll Learn

  • Recognize an exponential pattern that may become linear when the response is logged
  • Distinguish a regression of log response on the explanatory variable from a regression on the original response
  • Compare residual plots to assess whether a logarithmic transformation improves fit
  • Calculate a prediction on the log scale and convert it to the response’s original scale
  • Recognize when a logarithmic transformation does not adequately straighten a relationship

When a Straight Line Needs a Transformation

In “Combining Evidence of Linearity,” we used a scatterplot and residual plot to judge whether a linear model describes a relationship well. Sometimes the scatterplot bends, and the residual plot shows a curve: the straight-line model systematically misses the pattern. If the response changes by a roughly constant percentage as \(x\) changes, rather than by a constant amount, a logarithm of the response may turn the curve into a straighter pattern.

This tutorial focuses on transforming the response, not on changing the explanatory variable. We fit a regression line to \(\log(y)\) versus \(x\), inspect the residual plot for that model, and—when the transformed model is suitable—convert its prediction back to the original response scale. A transformation is a modeling choice to assess with evidence; it is not a guarantee that a relationship will become linear.

Definition: A logarithmic response transformation replaces each positive response \(y\) with its logarithm, such as \(\ln(y)\). If the plot of \(\ln(y)\) against \(x\) is approximately straight and its residual plot has no obvious systematic pattern, a linear model for \(\ln(y)\) may describe the relationship better than a line for \(y\).

How Logging Can Straighten an Exponential Pattern

An exponential relationship has a response that changes by a constant multiplicative factor for each one-unit increase in \(x\). One model for that pattern is \(y=A B^x\), where \(A>0\) and \(B>0\). Taking the natural logarithm of both sides changes multiplication and powers into addition and multiplication:

$$ \ln(y)=\ln(A)+x\ln(B) $$

This has the form of a line for the transformed response: the intercept is \(\ln(A)\), and the slope is \(\ln(B)\). The response must be positive to take its logarithm. The model’s slope is measured in log-response units per unit of \(x\), not in the original response units per unit of \(x\). In particular, it is not the constant amount by which the original response changes.

You can use a different logarithm base, such as \(\log_{10}\), provided you use that base consistently for the transformation and its inverse. The natural logarithm is convenient because the inverse operation is exponentiation by \(e\). The transformed regression equation predicts \(\ln(y)\); to express the prediction as an original-scale response, exponentiate that predicted value.

Formula: If the fitted line for the transformed response is \(\widehat{\ln(y)}=a+bx\), then the corresponding back-transformed fitted value is \(\hat{y}=e^{a+bx}\). If the transformation used \(\log_{10}\), use \(\hat{y}=10^{a+bx}\) instead.

The residuals for the transformed model are calculated on the log scale: observed \(\ln(y)\) minus predicted \(\widehat{\ln(y)}\). Judge that residual plot on its own scale. Look for residuals scattered around zero without a clear curve or other systematic pattern, as in the earlier tutorial “Reading a Residual Plot for Model Fit.” Compare the scatterplot and residual plot for the original-response line with the transformed model’s residual plot. An improved, more pattern-free residual plot is evidence that logging helped.

A back-transformed fitted value is a model-based prediction on the original scale. When residual scatter is present, it does not generally estimate the conditional arithmetic mean of \(y\) without an additional adjustment. Do not interpret the transformed line’s slope as an original-scale change in response units, and do not claim that a good fit proves a cause-and-effect relationship. As discussed in “Common Errors Linking Regression to Causation,” causal conclusions depend on study design.

A Practical Transformation-and-Prediction Process

1
Identify a possible exponential pattern.
Check whether the original scatterplot bends in a way that could reflect a roughly constant proportional change in the response. All response values used in a logarithm must be positive.
2
Fit a line to the logged response.
Use \(\ln(y)\) or another stated logarithm base as the response, and keep the original explanatory variable \(x\).
3
Assess the transformed model.
Compare the original scatterplot and residual evidence with the transformed model’s residual plot. An improved residual plot has less systematic curvature and no obvious new pattern.
4
Predict and return to original units.
Substitute the requested \(x\)-value into the transformed regression equation, then apply the inverse logarithm. Check whether the requested \(x\)-value is within the observed range and whether the prediction is sensible in context.

The residual plot is a diagnostic, not a promise that every prediction will be close. Even if logging improves the pattern, consider the amount of scatter, the observed \(x\)-range, and whether the predicted response is plausible. The tutorials “Interpolation Versus Extrapolation” and “Checking Whether a Prediction Is Reasonable” explain why those checks remain important after fitting a model.

Worked Examples: Transform, Diagnose, Predict

Worked Example: Logging a Curved Growth Pattern

Original AP-style question. In an invented demonstration, \(x\) is time in days and \(y\) is the amount of a growing culture, measured in milligrams. The observations are \((0,3),(1,6),(2,12),(3,24),(4,48)\). Compare the original-scale line with a line for \(\ln(y)\), describe the residual evidence, and use the transformed model to predict the amount at \(x=2.5\) days.

State. We want to see whether a logarithm makes the curved growth pattern more linear, then make an interpolated prediction in milligrams using the transformed model if it fits.

Plan. First find the least-squares line and residuals for \(y\) on \(x\). Then fit a line to \(\ln(y)\) on \(x\), compare its residuals with the original model’s residuals, and back-transform the prediction at \(x=2.5\). All five responses are positive, so their logarithms are defined.

Do. For the original responses, \(\bar{x}=2\) and \(\bar{y}=18.6\). The centered \(x\)-values are \(-2,-1,0,1,2\), so \(S_{xx}=10\). The cross-product sum is \(S_{xy}=(-2)(3-18.6)+(-1)(6-18.6)+0(12-18.6)+1(24-18.6)+2(48-18.6)=108\). Thus the slope is \(108/10=10.8\), and the intercept is \(18.6-10.8(2)=-3\). The original-scale fitted line is \(\hat{y}=-3+10.8x\).

$$ \begin{array}{c|ccccc} x & 0 & 1 & 2 & 3 & 4\\ \text{Original-scale residual} & 6 & -1.8 & -6.6 & -5.4 & 7.8 \end{array} $$

For example, at \(x=0\) the predicted response is \(-3+10.8(0)=-3\), so the residual is \(3-(-3)=6\) milligrams. At \(x=2\), the prediction is \(18.6\) and the residual is \(12-18.6=-6.6\) milligrams. The residuals tend to be positive at the ends and negative in the middle, a curved pattern consistent with the bend in the original scatterplot.

Now take natural logarithms. The transformed responses are \(\ln(3),\ln(6),\ln(12),\ln(24),\ln(48)\), or approximately \(1.0986,1.7918,2.4849,3.1781,3.8712\). Each is \(\ln(3)+x\ln(2)\). Therefore the transformed data lie on the line \(\widehat{\ln(y)}=1.0986+0.6931x\), with residuals zero apart from rounding. Its residual plot has no curved pattern, so logging improves the linear fit for these data.

At \(x=2.5\), the predicted log response is \(1.0986+0.6931(2.5)=2.8314\), using rounded coefficients. Back-transforming gives \(e^{2.8314}\approx16.97\) milligrams. Equivalently, the exact pattern gives \(3(2^{2.5})\approx16.97\) milligrams. Since \(2.5\) is between the observed \(x\)-values 0 and 4, this is interpolation.

Conclude. The original-scale residuals show curvature, while the transformed residuals are approximately zero and pattern-free. The log transformation provides a much straighter model for this invented data set, which predicts about 16.97 milligrams at 2.5 days. The exact fit here is especially tidy; real data generally have some residual scatter.

Worked Example: Transforming a Decreasing Exponential Pattern

Original AP-style question. In an invented equipment exercise, \(x\) is elapsed time in hours and \(y\) is the remaining amount of a substance in grams. The observations are \((0,81),(1,27),(2,9),(3,3),(4,1)\). Compare the residual pattern of a line for \(y\) with a line for \(\log_{10}(y)\), then predict the amount remaining after 2.5 hours.

State. We will assess whether logging the positive response makes the decreasing curved relationship more suitable for a linear model, then calculate and interpret the back-transformed prediction.

Plan. Fit a line to each response scale and compare residual patterns. For the logged model, use base 10 throughout and convert the prediction using the inverse operation, raising 10 to the predicted log value.

Do. For the original responses, \(\bar{x}=2\), \(\bar{y}=24.2\), and \(S_{xx}=10\). The cross-product sum is \(S_{xy}=(-2)(81-24.2)+(-1)(27-24.2)+0(9-24.2)+1(3-24.2)+2(1-24.2)=-184\). Thus \(b=-184/10=-18.4\) and \(a=24.2-(-18.4)(2)=61\). The original-scale line is \(\hat{y}=61-18.4x\). Its fitted values at \(x=0,1,2,3,4\) are \(61,42.6,24.2,5.8,-12.6\), giving residuals \(20,-15.6,-15.2,-2.8,13.6\) grams. This pattern is not a random band around zero; the residuals are positive at both ends and mostly negative in the middle.

The base-10 logarithms of the responses are approximately \(1.908485,1.431364,0.954243,0.477121,0\). They follow the line \(\log_{10}(y)=\log_{10}(81)-x\log_{10}(3)\), so the transformed regression equation is \(\widehat{\log_{10}(y)}=1.908485-0.477121x\). Its residuals are zero apart from rounding, an improvement over the curved original-scale residual pattern.

At \(x=2.5\), the predicted log response is \(1.908485-0.477121(2.5)=0.715682\). The prediction in grams is \(10^{0.715682}\approx5.196\) grams. This can be checked from the exact pattern: \(81(1/3)^{2.5}\approx5.196\). The requested time lies between 0 and 4 hours, so the prediction is an interpolation.

Conclude. The transformed residuals show a straighter pattern than the original-scale residuals, supporting the logarithmic model for these data. The back-transformed model predicts about 5.196 grams remaining after 2.5 hours. The transformed slope describes change in \(\log_{10}(y)\) per hour; it does not mean that the amount decreases by 0.477121 grams each hour.

Worked Example: Logging Does Not Straighten Every Curve

Original AP-style question. In an invented geometry exercise, \(x\) is a measured length and \(y\) is a calculated area in square units. The data are \((0,1),(1,4),(2,9),(3,16),(4,25)\). Fit a line to \(\ln(y)\), inspect its residuals, and explain whether this transformation supports using a straight line to predict area.

State. The original relationship is curved, but that alone does not show that logging the response will fix it. We will check whether the transformed residuals are pattern-free.

Plan. Calculate the least-squares line for \(\ln(y)\) on \(x\), then find residuals in order of \(x\). A clear remaining pattern means that the log transformation has not made a linear model adequate.

Do. The logged responses are \(0,\ln(4),\ln(9),\ln(16),\ln(25)\), approximately \(0,1.386294,2.197225,2.772589,3.218876\). Their mean is \(1.914997\). As before, \(\bar{x}=2\) and \(S_{xx}=10\). The cross-product sum is

$$ S_{xy}=(-2)(0-1.914997)+(-1)(1.386294-1.914997) +(1)(2.772589-1.914997)+2(3.218876-1.914997) \approx7.824046 $$

The slope is \(7.824046/10\approx0.782405\), and the intercept is \(1.914997-0.782405(2)\approx0.350187\). Thus \(\widehat{\ln(y)}=0.350187+0.782405x\). The residuals, calculated as observed \(\ln(y)\) minus predicted \(\widehat{\ln(y)}\), are approximately \(-0.350188,0.253702,0.282228,0.075187,-0.260930\). They are negative at both ends and positive in the middle, which is a curved pattern.

For illustration, at \(x=2.5\) the fitted log response is \(0.350187+0.782405(2.5)\approx2.306199\), so the back-transformed fitted value is \(e^{2.306199}\approx10.04\) square units. The requested \(x\) is within the observed range, but the transformed residual curve warns that this linear model does not capture the pattern well. Therefore, that calculation should not be presented as a dependable prediction from an adequate model.

Conclude. Logging the response has not removed the systematic curve in the residuals. This example shows why a transformation must be assessed by the resulting plots rather than assumed to work just because it is a logarithm.

Common Mistakes and AP Exam Tip

  • Logging a nonpositive response. The logarithm is defined only for positive values in this setting. Check the response values before transforming; do not silently take the logarithm of zero or a negative number.
  • Confusing which variable was transformed. Here the model is a regression of \(\ln(y)\) on \(x\). State clearly that the response was transformed and that \(x\) remains on its original scale.
  • Interpreting the slope in original response units. In \(\widehat{\ln(y)}=a+bx\), \(b\) is in log-response units per \(x\)-unit. It is not a change of \(b\) milligrams, grams, or other original units for each one-unit increase in \(x\).
  • Stopping at the predicted logarithm. A value such as \(2.8314\) in a natural-log model is a predicted log response, not the response in its original units. Apply the inverse transformation and label the resulting units.
  • Calling a transformed model adequate because its scatterplot looks straighter. Check the transformed model’s residual plot for remaining curves, changing spread, or other systematic patterns. A transformation may help without solving every fit problem.
  • Forgetting the prediction’s scope. Check the observed range of \(x\) and the plausibility of the back-transformed prediction, as in “Reliability Within the Range of Data” and “Impossible or Nonsensical Predictions.” A better residual plot does not eliminate extrapolation concerns.

A strong AP response identifies the transformation, names the model being assessed, and describes the residual evidence specifically. For example: “After regressing \(\ln(y)\) on \(x\), the residuals are scattered around zero without the curved pattern seen for the original response, so the logarithmic transformation improves the linear fit. The predicted log response must be exponentiated to report a prediction in the original units.” If a curve remains, say so rather than claiming that the transformation worked.

Key takeaway: Logging a positive response can straighten an exponential relationship because it converts a multiplicative pattern into a linear one. Use the transformed residual plot to judge whether the model improved, and apply the inverse logarithm to express a prediction in the original response units.

Check Your Understanding

For each question, identify the response scale and describe the evidence or calculation needed.

  1. Why might a response that grows by roughly the same percentage for each one-unit increase in \(x\) become more linear when you plot \(\ln(y)\) against \(x\)?
  2. A fitted model is \(\widehat{\ln(y)}=0.4+0.2x\). Write the corresponding back-transformed fitted value at \(x=3\).
  3. What residual-plot evidence would support the claim that a logarithmic transformation improved a curved relationship?
  4. A transformed residual plot still has residuals negative at both ends and positive in the middle. What does that pattern suggest about the model?
  5. Why is the slope in a regression of \(\ln(y)\) on \(x\) not an amount of \(y\)-units gained for each one-unit increase in \(x\)?