From Paired Data to a Regression Equation
In “Finding the Line Through the Means,” you used the fact that a least-squares regression line passes through \((\bar{x},\bar{y})\). Now you can use a TI-84 to calculate the line directly from paired data. The calculator returns two important values: the intercept \(a\) and the slope \(b\).
The command used here is LinReg(a+bx). Its name tells you the form of the equation: intercept \(a\), plus slope \(b\) times predictor \(x\). The calculator uses \(x\) for the first list you specify and \(y\) for the second. Those letters describe the lists’ roles in the calculation; in your final answer, use variable names that identify the predictor and response in context.
The order of the lists matters. Put the predictor values in the first list and the corresponding response values in the second. Each row is one case, so the \(x\)-value and \(y\)-value on that row must belong together. The calculator does not know which variable you intend to explain; you establish that by how you enter the lists.
Running LinReg(a+bx) on a TI-84
First enter the paired data. Press STAT, choose EDIT, and enter the predictor values in \(L1\) and the matching response values in \(L2\). Check that the lists have the same number of values and that each pair is in the intended row.
Next, open the STAT menu, go to CALC, and select LinReg(a+bx). Supply \(L1,L2\) in that order and run the command. Menu numbering can differ across TI-84 models, so select the command by its name rather than relying only on an option number. The output includes \(a\) and \(b\). Read those labels carefully: in this command, \(a\) is the intercept and \(b\) is the slope.
Enter the predictor in \(L1\), the response in \(L2\), and keep every pair on the same row.
Select LinReg(a+bx) from STAT, CALC, and use \(L1,L2\) as the lists.
Record the output beside \(a\) as the intercept and the output beside \(b\) as the slope.
Use the response name with a hat on the left and the predictor name on the right, preserving the intercept and slope.
For example, if \(P\) is practice time and \(S\) is a test score, the calculator’s generic equation \(\hat{y}=a+bx\) becomes \(\hat{S}=a+bP\). The hat indicates a predicted score, not an observed score. Use parentheses around a negative coefficient when needed to make the arithmetic clear.
A calculator can fit a line whether or not a linear model is sensible. Before using the equation to describe the data, look at the scatterplot for a roughly straight pattern and check for unusual observations, as in “Strong Correlation Does Not Mean a Good Model.” LinReg calculates the line; it does not decide whether a line is a useful summary.
Worked Examples
Worked Example: Practice Time and Quiz Score
A fictional teacher records \(P\), hours spent practicing, and \(S\), quiz score in points, for five students. The paired observations are:
| Practice time \(P\) (hours) | Quiz score \(S\) (points) |
|---|---|
| 1 | 52 |
| 2 | 55 |
| 3 | 59 |
| 4 | 60 |
| 5 | 64 |
Enter \(1,2,3,4,5\) in \(L1\) and \(52,55,59,60,64\) in \(L2\). Run LinReg(a+bx) \(L1,L2\). The calculator reports \(a=49.3\) and \(b=2.9\). The first list contains practice time, so \(x=P\); the second contains quiz score, so \(y=S\).
Thus the fitted equation, with the variables named, is:
Here, \(\hat{S}\) is the predicted quiz score in points for a student who practiced \(P\) hours. The slope is \(2.9\) points per hour, and the intercept is \(49.3\) points. As a check on the calculator output, the sample means are \(\bar{P}=3\) hours and \(\bar{S}=58\) points. The slope calculated from the paired data is:
Then the mean-point relationship from “Finding the Line Through the Means” gives \(a=\bar{S}-b\bar{P}=58-(2.9)(3)=49.3\). Substituting \(P=3\) in the equation gives \(\hat{S}=49.3+(2.9)(3)=58\), as expected. These checks confirm that the coefficients match the data and their labels.
Worked Example: Equipment Age and Efficiency Rating
A fictional maintenance team records \(A\), equipment age in years, and \(E\), an efficiency rating in points, for five machines. The paired values are:
| Age \(A\) (years) | Efficiency \(E\) (points) |
|---|---|
| 2 | 31 |
| 4 | 28 |
| 6 | 26 |
| 8 | 23 |
| 10 | 22 |
Enter age in \(L1\) and efficiency in \(L2\), then run LinReg(a+bx) \(L1,L2\). The calculator reports \(a=32.9\) and \(b=-1.15\). Since \(A\) is the predictor and \(E\) is the response, the regression equation is:
The negative slope is part of the reported \(b\), not the intercept. To verify it, \(\bar{A}=6\) years and \(\bar{E}=26\) points. The sum of cross-products is \(-46\), and the sum of squared deviations in age is \(40\), so \(b=-46/40=-1.15\). Then \(a=\bar{E}-b\bar{A}=26-(-1.15)(6)=26+6.9=32.9\). The equation predicts \(26\) points at the mean age of six years: \(32.9-1.15(6)=26\).
In context, the fitted line predicts an efficiency rating \(1.15\) points lower for each additional year of age. The intercept is the predicted rating at age zero. Whether that prediction is useful depends on the machines and age range represented in the data.
Worked Example: Temperature and Plant Growth
A fictional greenhouse trial records \(T\), temperature in degrees Celsius, and \(G\), plant growth in centimeters, for four plots. The paired observations are:
| Temperature \(T\) (degrees Celsius) | Growth \(G\) (centimeters) |
|---|---|
| 10 | 4 |
| 20 | 7 |
| 30 | 8 |
| 40 | 11 |
Enter the temperatures in \(L1\) and the matching growth values in \(L2\). LinReg(a+bx) \(L1,L2\) reports \(a=2\) and \(b=0.22\). The requested equation with meaningful variable names is:
Check the slope using the data summaries. The means are \(\bar{T}=25\) degrees Celsius and \(\bar{G}=7.5\) centimeters. The sum of cross-products is \(110\), and the sum of squared temperature deviations is \(500\). Therefore \(b=110/500=0.22\) centimeters per degree Celsius. The intercept is \(a=7.5-(0.22)(25)=7.5-5.5=2\) centimeters.
The equation predicts growth in centimeters, so \(G\), not \(T\), belongs on the left. It predicts \(0.22\) centimeters more growth for each additional degree Celsius in temperature. Reversing the letters would describe a different regression and would not answer the question of predicting growth from temperature.
Keep the Variable Roles and Coefficients Straight
The command’s form matters. In LinReg(a+bx), \(a\) is added on its own and \(b\) multiplies \(x\). Do not assume that the letter \(a\) always means slope or that \(b\) always means intercept: use the output labels for the specific command you selected. A different command form can label coefficients differently, so match your equation to LinReg(a+bx).
Likewise, list order determines the roles. If you put the response in \(L1\) and the predictor in \(L2\), you are asking for a line that predicts the original predictor from the response. In general, that is not the same line with the variables exchanged. Use the order required by the question: predictor first, response second.
The calculator’s \(x\) and \(y\) are placeholders for the values in your lists. A final answer that only says \(\hat{y}=49.3+2.9x\) may be ambiguous when the question gives named variables. State what \(x\) and \(y\) represent, or replace them with the variable names, such as \(\hat{S}=49.3+2.9P\). Keep the hat on the predicted response.
Common Mistakes and AP Exam Tips
- Swapping the lists. Enter the explanatory variable in \(L1\) and the response variable in \(L2\). Confirm the list headings and the order in the command before running it.
- Interchanging \(a\) and \(b\). For LinReg(a+bx), \(a\) is the intercept and \(b\) is the slope. Check the command form and output labels rather than relying on a memorized letter rule.
- Dropping a negative sign. If the output gives \(b=-1.15\), write the equation with a negative slope, such as \(\hat{E}=32.9-1.15A\).
- Leaving variables unnamed. The generic calculator equation is not always a complete contextual answer. Identify the predictor and response, and write the predicted response with a hat.
- Pairing the wrong observations. Keep each predictor and response on the same row. A correct calculator command cannot repair mismatched pairs.
- Trusting a line without checking the plot. Inspect the scatterplot for a roughly linear pattern and unusual points. The calculator will return coefficients even when a straight-line model is not a sensible summary.
- Rounding too early. Use the calculator’s displayed coefficients at suitable precision, and avoid altering the slope or intercept during intermediate checks. Report rounded values consistently.
A clear AP-style response names the lists’ roles, identifies the calculator output, and writes the equation with the contextual variables. For example: “With practice time \(P\) in \(L1\) and quiz score \(S\) in \(L2\), LinReg(a+bx) gives \(a=49.3\) and \(b=2.9\). The fitted line is \(\hat{S}=49.3+2.9P\), where \(\hat{S}\) is predicted quiz score in points.”
Check Your Understanding
Use the calculator’s LinReg(a+bx) convention and keep the predictor and response roles clear.
- What should go in \(L1\) and \(L2\) when predicting daily water use from household size?
- A LinReg(a+bx) output gives \(a=18.6\) and \(b=-0.4\). Write the generic regression equation.
- A study uses \(D\) for delivery distance in kilometers and \(C\) for delivery cost in dollars. The calculator reports \(a=3.50\) and \(b=0.80\), with distance in \(L1\) and cost in \(L2\). Write an equation using the named variables.
- Why can’t you simply swap \(L1\) and \(L2\) and expect the same regression line with the variable names reversed?
- What should you inspect before treating a calculator-generated regression equation as a useful summary of the data?