Tutorials › AP Statistics › Writing the Equation With Variable Names

Linear regression models · Tutorial 854 of 1000

Writing the Equation With Variable Names

Practice writing fitted regression equations with names that show clearly which variable is the predictor and which response is being predicted.

Intermediate 9 min read

What You'll Learn

  • Replace \(x\) with a descriptive name for the predictor in a regression equation.
  • Mark the predicted response with a hat on its descriptive name.
  • Keep the intercept as the constant term and the slope multiplying the predictor.
  • Use variable names and units to make an equation easier to read in context.
  • Check that a written equation preserves the correct coefficient roles and signs.

Make the Regression Equation Say What It Predicts

In “Reading the Equation of a Regression Line,” the fitted line was written as \(\hat{y}=a+bx\). That form is compact, but it uses symbols that do not identify the variables on their own. When a problem gives variables such as “daily screen time” and “battery drain,” writing the equation with those names can make the predictor and the predicted response much easier to recognize.

The task is a change in notation, not a different regression model. The predictor still takes the place of \(x\), the predicted response still takes the place of \(\hat{y}\), the intercept remains the constant term, and the slope still multiplies the predictor. This is the same line discussed in “Finding the Equation With LinReg(a+bx)” and “Reading Regression Computer Output,” now written with variable names in place of generic symbols.

Definition: A regression equation with descriptive variable names writes the predicted response using a name for that response and writes the predictor by its name. In \(\hat{y}=a+bx\), the descriptive-name version places a hat on the response name and multiplies the slope by the predictor name.

For example, if a line predicts travel time from route distance, you could write \(\widehat{\text{travel time}}=a+b(\text{route distance})\). The hat indicates that the left side is a predicted value, not an observed travel time. The name in parentheses identifies the predictor whose value is used to obtain that prediction.

The parentheses are useful when a variable name contains multiple words. They show that the slope is multiplied by the predictor value as a whole. They are not essential if the notation is otherwise clear, but they help avoid reading adjacent words or symbols as part of the coefficient.

A Consistent Translation from Generic Symbols

Before substituting coefficients, name the variables in their roles. The response is the variable the line predicts; the predictor is the explanatory variable used as input. These roles were central in “Reading Regression Computer Output.” A variable’s position in a written equation should reflect its role, not simply the order in which the problem mentions it.

1
Name the response and predictor.
Write down which quantity is being predicted and which quantity supplies the input to the line.
2
Replace \(\hat{y}\) with the predicted response name.
Put a hat over the response name to show that the equation gives a predicted value.
3
Replace \(x\) with the predictor name.
Place the predictor name beside the slope, using parentheses if that makes the multiplication clearer.
4
Insert the coefficients in their roles.
Write the intercept as the constant and the slope as the number multiplying the predictor. Preserve the slope’s sign.

In symbols, the translation is:

$$ \hat{y}=a+bx \quad\longrightarrow\quad \widehat{\text{predicted response}}=a+b(\text{predictor}) $$

The words “predicted response” and “predictor” in this general pattern are placeholders. Replace them with the actual variable names from the situation. For instance, if the response is weekly water use and the predictor is the number of hot days, use those names—not the generic placeholders—in the final equation.

Keep the hat on the response name, not on the predictor name. The predictor is an input value; the left side is the predicted response. This distinction also keeps the equation separate from an observed pair of measurements. An observed response is not necessarily equal to its predicted value; as covered in “Predicted Change Versus Actual Change,” observations can differ from the line’s predictions.

Worked Examples

Worked Example: Outdoor Temperature and Greenhouse Water Use

A fictional greenhouse uses outdoor temperature to predict daily water use. Temperature is measured in degrees Celsius, and water use is measured in liters. A software table reports an intercept estimate of \(18.4\) and a slope estimate of \(2.7\). Write the fitted equation using descriptive variable names.

The response is daily water use, because that is the quantity being predicted. The predictor is outdoor temperature. In the generic form, the intercept is \(a=18.4\) and the slope is \(b=2.7\). Thus, \(\hat{y}\) becomes \(\widehat{\text{daily water use}}\), and \(x\) becomes \(\text{outdoor temperature}\).

$$ \widehat{\text{daily water use}} =18.4+2.7(\text{outdoor temperature}) $$

This equation predicts daily water use in liters from outdoor temperature in degrees Celsius. The name on the left is marked with a hat because it is a predicted response. The predictor name appears multiplied by the slope, and the intercept remains the constant term. The descriptive equation represents the same fitted line that could be written as \(\hat{y}=18.4+2.7x\); the names clarify what the symbols stand for.

Worked Example: Cycling Distance and Battery Charge Used

A fictional technology club models the percentage points of battery charge used by a bicycle computer from the distance ridden, measured in kilometers. The intercept is \(5.6\), and the slope is \(-0.18\). Write the regression equation with descriptive names.

First identify the roles: battery charge used is the response, and cycling distance is the predictor. Substituting into the generic form gives a predicted-response name on the left, the intercept \(5.6\) as the constant, and the slope \(-0.18\) multiplying the predictor name. The negative sign must remain attached to the slope.

$$ \widehat{\text{battery charge used}} =5.6+(-0.18)(\text{cycling distance}) $$

It is also standard to simplify the addition of a negative term:

$$ \widehat{\text{battery charge used}} =5.6-0.18(\text{cycling distance}) $$

Both equations preserve the same intercept and slope. In either version, the left side is the predicted response, while cycling distance is the predictor. The measurement units are percentage points for the response and kilometers for the predictor. The slope’s units are percentage points per kilometer, consistent with the rate-of-change convention in “Slope Units and Rates of Change.”

Worked Example: Number of Open Counters and Store Wait Time

A fictional store examines whether the number of open checkout counters can be used to predict customer wait time in minutes. A computer output gives these coefficient estimates:

TermEstimate
Number of open counters-1.35
Intercept11.2

The response is customer wait time, and the predictor is number of open counters. The row labelled “Intercept” supplies \(a=11.2\). The predictor row supplies \(b=-1.35\). The order of the rows does not change these roles, as emphasized in “Reading Regression Computer Output.”

Replace \(\hat{y}\) with the name of the predicted response and \(x\) with the name of the predictor. Then put the intercept in the constant position and the negative slope beside the predictor:

$$ \widehat{\text{customer wait time}} =11.2+(-1.35)(\text{number of open counters}) $$

Writing the negative term with subtraction gives the same equation:

$$ \widehat{\text{customer wait time}} =11.2-1.35(\text{number of open counters}) $$

The first form makes the negative coefficient explicit; the second is more compact. Both use descriptive names and preserve the correct coefficient roles. Customer wait time is the predicted response in minutes, and the number of open counters is the predictor. This example is about writing the equation; any interpretation of its intercept or slope should be handled separately and in context.

Choose Names That Make the Roles Clear

A descriptive name should be specific enough to identify the measured quantity. “Time” may be unclear if the situation includes several kinds of time. “Customer wait time” is more informative. Similarly, if a problem distinguishes distance ridden from distance remaining, use the name that exactly matches the predictor used to fit the line.

Include units in the surrounding explanation when the variable name alone does not show them. For example, you might define the response as “daily water use, in liters” and the predictor as “outdoor temperature, in degrees Celsius,” then write the equation using the shorter names. Units are not a substitute for identifying the variables, but they help make the equation unambiguous.

Long names can make equations visually crowded. It is acceptable to define short labels first, as long as the link to the real variables is clear. For example, define \(W\) as daily water use in liters and \(T\) as outdoor temperature in degrees Celsius, then write \(\widehat{W}=18.4+2.7T\). This still uses letters, but the definitions make their meanings explicit. If the instruction specifically asks for descriptive variable names, show the names in the final equation rather than relying only on the abbreviations.

Key takeaway: To write a regression equation with descriptive names, put a hat on the name of the predicted response, use the predictor’s name beside the slope, and leave the intercept as the constant. The result is the same fitted line with its variable roles made explicit.

Common Mistakes and AP Exam Tips

  • Putting the predictor on the left. The left side is the predicted response, not the predictor. Identify the variable the model predicts before writing the equation.
  • Leaving off the hat. A hat distinguishes a predicted response from an observed response. Write \(\widehat{\text{wait time}}\), for example, when the equation gives predicted wait time.
  • Putting a hat on the predictor. The predictor is the input to the line. Keep the hat on the predicted response name, not on the name in parentheses.
  • Writing the slope as a constant. The slope multiplies the predictor. An equation such as \(\widehat{\text{water use}}=18.4+2.7\) leaves out the predictor and is not the fitted line.
  • Reversing the intercept and slope. The intercept is the constant term; the slope is attached to the predictor. Use the coefficient roles established in “Reading Regression Computer Output,” not the order of terms in a table.
  • Dropping or changing the sign. Copy a negative slope as negative. Writing \(+1.35(\text{number of open counters})\) instead of \(-1.35(\text{number of open counters})\) changes the fitted equation.
  • Using names that do not match the question. A generic label such as “amount” can leave the response unclear. Name the actual measured quantity and, where useful, state its units.

A full-credit response makes the roles visible. For example: “Customer wait time is the response, and number of open counters is the predictor. Therefore, the fitted equation is \(\widehat{\text{customer wait time}}=11.2-1.35(\text{number of open counters})\), with wait time measured in minutes.” That statement identifies the variables and writes the equation without requiring the reader to guess what \(x\) and \(y\) represent.

Check Your Understanding

For each item, identify the response and predictor and write an equation using descriptive variable names.

  1. A line predicts commute time from commute distance. The intercept is \(8.3\), and the slope is \(1.6\). Write the equation using commute time and commute distance.
  2. A model predicts weekly printing cost from the number of pages printed. Its intercept is \(12\), and its slope is \(0.04\). Which variable belongs on the left side with a hat?
  3. A regression line predicts the number of seedlings surviving from hours of daily sunlight. The intercept is \(42\), and the slope is \(-1.8\). Write the equation and preserve the sign.
  4. In a descriptive-name equation, where should the intercept appear, and what should multiply the slope?
  5. Why does a hat belong on the response name rather than the predictor name?