Tutorials › AP Statistics › Finding r With LinReg on the Calculator

Correlation · Tutorial 825 of 1000

Finding r With LinReg on the Calculator

Use the TI-84’s LinReg(a+bx) command and diagnostics to find the correlation \(r\) and coefficient of determination \(r^2\).

Intermediate 9 min read

What You'll Learn

  • Turn on calculator diagnostics so LinReg(a+bx) displays \(r\) and \(r^2\)
  • Enter paired quantitative data into L1 and L2 and run the linear regression command
  • Tell the calculator’s intercept, slope, \(r^2\), and \(r\) apart
  • Check calculator output using deviations from the means
  • Describe \(r\) and \(r^2\) in context without making a causal claim

Let the Calculator Report \(r\)

In “Calculating \(r\) From Standardized Values,” you found the correlation by standardizing paired observations and combining their z-scores. For a larger data set, a calculator can do that arithmetic quickly. The TI-84’s LinReg(a+bx) command fits a least-squares regression line and, when diagnostics are on, reports both \(r\) and \(r^2\).

The command’s name describes the line it fits: \(y=a+bx\), where \(a\) is the y-intercept and \(b\) is the slope. These are regression quantities, not correlation. The diagnostic values appear separately. In particular, \(r\) is the correlation between the explanatory and response variables; \(r^2\) is the square of that correlation and is also called the coefficient of determination.

Definition: In the LinReg(a+bx) output, \(a\) is the intercept, \(b\) is the slope, \(r^2\) is the coefficient of determination, and \(r\) is the correlation coefficient. With diagnostics enabled, a TI-84 typically displays \(r^2\) before \(r\). Read the labels rather than identifying values by their position alone.

For \(r\) to be defined, both variables must vary. As in “What the Correlation Coefficient Measures,” it summarizes the direction and strength of a linear association between two quantitative variables. Before relying on it, inspect the scatterplot: as discussed in “Matching Correlations to Scatterplots,” a curved pattern or an influential unusual point can make a single correlation misleading.

Turn Diagnostics On and Run LinReg(a+bx)

Some TI-84 calculators have diagnostics turned off initially. In that case, a LinReg command can show \(a\) and \(b\) but omit \(r^2\) and \(r\). Turn diagnostics on once, then run the regression again. The setting generally remains on until it is changed.

1
Enter the paired data.
Press STAT, choose EDIT, and enter the explanatory-variable values in L1 and their matching response-variable values in L2. Each row must contain a pair from the same individual or case.
2
Enable diagnostics.
Press 2nd, then 0 to open the catalog. Scroll to DiagnosticOn, select it, and press ENTER. Press ENTER again to execute it. The calculator should display “Done.”
3
Choose LinReg(a+bx).
Press STAT, open the CALC menu, and select LinReg(a+bx). The menu entry may be farther down the list, so scroll to the command rather than relying on a menu number.
4
Supply the lists and run the command.
Enter L1, a comma, and L2, then select Calculate or press ENTER. The command-line form is LinReg(a+bx) L1,L2.
5
Read the labelled output.
Record \(a\), \(b\), \(r^2\), and \(r\). If \(r\) and \(r^2\) are missing, turn on DiagnosticOn and rerun the command.

List order matters. Put the explanatory variable in L1 and the response variable in L2, following the convention from “Making a Scatterplot on the TI-84.” Swapping the lists does not change \(r\) or \(r^2\), but it changes the regression equation: the slope and intercept then describe a different response variable.

Formula: For a linear regression with an intercept, \(r^2\) is the square of \(r\). Therefore \(r=\sqrt{r^2}\) when the association is positive and \(r=-\sqrt{r^2}\) when it is negative. The sign must come from the direction of the association; \(r^2\) alone has no sign.

Worked Example: A Positive Association

Worked Example: A Positive Association

A fictional class records the number of practice sessions \(x\) and a skills-check score \(y\), in points, for five participants. The values are invented for this example. Enter \(x\) in L1 and \(y\) in L2.

ParticipantPractice sessions \(x\)Skills-check score \(y\), points
112
224
333
445
556

Calculator procedure. After entering the lists and confirming that DiagnosticOn is enabled, run LinReg(a+bx) L1,L2. The calculator reports \(a=1.3\), \(b=0.9\), \(r^2=0.81\), and \(r=0.9\). The fitted line is \(\hat{y}=1.3+0.9x\), but the correlation is the separately labelled \(r\), not the slope \(b\).

Check the correlation. The means are \(\bar{x}=3\) sessions and \(\bar{y}=4\) points. The sums of squared deviations are \(\sum(x-\bar{x})^2=10\) and \(\sum(y-\bar{y})^2=10\). The sum of paired deviation products is \(9\): the products are \(4,0,0,1,4\). Thus the deviation formula gives:

$$ r=\frac{\sum(x-\bar{x})(y-\bar{y})} {\sqrt{\sum(x-\bar{x})^2\sum(y-\bar{y})^2}} =\frac{9}{\sqrt{10(10)}}=0.9 $$

Squaring gives \(r^2=(0.9)^2=0.81\), which agrees with the calculator. In context, these five participants show a positive linear association between practice sessions and skills-check score: participants with more sessions tended to have higher scores. The fitted line’s slope of \(0.9\) points per session is not the correlation; the correlation \(0.9\) is unitless.

Worked Example: A Negative Association

Worked Example: A Negative Association

A fictional recreation program records the number of days participants used a route-planning app \(x\) and the number of missed turns \(y\) on a practice route. Enter app-use days in L1 and missed turns in L2.

ParticipantApp-use days \(x\)Missed turns \(y\)
1110
228
339
446
557

Calculator output. With diagnostics on, LinReg(a+bx) L1,L2 reports \(a=10.4\), \(b=-0.8\), \(r^2=0.64\), and \(r=-0.8\). The slope is negative, which agrees with the downward direction of the association, but its units are missed turns per app-use day. The correlation is unitless.

Check the sign and value. Here \(\bar{x}=3\) days and \(\bar{y}=8\) missed turns. The x-deviations are \(-2,-1,0,1,2\), and the y-deviations are \(2,0,1,-2,-1\). Their paired products sum to \(-8\), while each variable’s squared deviations sum to \(10\). Therefore:

$$ r=\frac{-8}{\sqrt{10(10)}}=-0.8 \qquad\text{and}\qquad r^2=(-0.8)^2=0.64 $$

The negative value of \(r\) matches the pattern: participants with more app-use days tended to have fewer missed turns in these observations. The positive value of \(r^2\) does not mean the association is positive. Squaring removes the sign, so use \(r\), together with the scatterplot, to describe direction.

Worked Example: Read Both Diagnostics Correctly

Worked Example: Read Both Diagnostics Correctly

A fictional technology club records weekly practice time \(x\), in hours, and a task score \(y\), in points. A student runs LinReg(a+bx), but the first output shows only \(a\) and \(b\). The student turns on DiagnosticOn as described above and runs the command again.

MemberPractice time \(x\), hoursTask score \(y\), points
113
224
335
444
556

Read the output. The calculator reports \(a=2.6\), \(b=0.6\), \(r^2\approx0.6923\), and \(r\approx0.8321\), with values rounded to four decimal places where needed. The slope and intercept give the fitted line \(\hat{y}=2.6+0.6x\); \(r\) and \(r^2\) summarize the linear association and its fit.

Verify the diagnostics. The means are \(\bar{x}=3\) hours and \(\bar{y}=4.4\) points. The x-deviations have squared sum \(10\). The y-deviations are \(-1.4,-0.4,0.6,-0.4,1.6\), whose squared deviations sum to \(5.2\). The paired deviation products are \(2.8,0.4,0,-0.4,3.2\), with sum \(6\). So:

$$ r=\frac{6}{\sqrt{10(5.2)}}=\frac{6}{\sqrt{52}} \approx0.8321 $$

Squaring the unrounded value gives \(r^2=36/52=9/13\approx0.6923\). The calculator’s two diagnostic values agree. In context, the five club members show a fairly strong positive linear association between weekly practice time and task score. About \(69.23\%\) of the variation in task scores among these members is accounted for by the least-squares regression of score on practice time. This is a description of the fitted linear relationship, not evidence that practice time caused the scores.

What \(r^2\) Adds

The correlation \(r\) communicates both direction and strength of a linear association. The coefficient of determination \(r^2\) communicates how much of the variation in the response variable is accounted for by its linear regression on the explanatory variable. Because \(r^2\) is a proportion, it ranges from \(0\) to \(1\), and it has no direction or units.

For example, \(r^2=0.64\) means that \(64\%\) of the variation in the response values in the data is accounted for by the least-squares regression line using the explanatory variable. It does not mean that \(64\%\) of individuals are predicted correctly, that the model is accurate for every observation, or that the explanatory variable caused the response. Interpret the percentage in context and keep the response variable clear.

Key takeaway: On a TI-84 with diagnostics on, read \(r\) for the signed correlation and \(r^2\) for the proportion of response variation accounted for by the linear regression. Do not confuse either value with the slope or intercept.

Common Mistakes and AP Exam Tips

  • Assuming \(r\) always appears. If diagnostics are off, the calculator may omit \(r\) and \(r^2\). Turn on DiagnosticOn and rerun the regression; do not guess \(r\) from the slope.
  • Mixing up \(r\) and \(r^2\). The square \(r^2\) is nonnegative and gives no direction. Use the labelled \(r\) to report whether the linear association is positive or negative.
  • Calling the slope a correlation. The slope \(b\) has units, while \(r\) is unitless and always lies between \(-1\) and \(1\). A slope can be much larger than 1 or less than \(-1\).
  • Swapping the lists accidentally. Confirm that each L1 entry is paired with the matching L2 entry. Swapping explanatory and response variables changes the regression line, even though \(r\) remains the same.
  • Interpreting \(r^2\) as causation or prediction accuracy. State the percent of variation in the response accounted for by the linear regression on the explanatory variable. Do not claim that the relationship is causal or that the same percentage of individual predictions is correct.
  • Reporting a correlation without checking the scatterplot. A curved pattern or unusual point can make \(r\) an incomplete summary. Use the scatterplot and DUFS ideas from earlier tutorials to check form and unusual features.
  • Rounding too early or misreading the display. Keep the calculator’s full precision for checks, and use the output labels. If you square a rounded display of \(r\), the result may differ slightly from the displayed \(r^2\).

For full-credit communication, identify the response and explanatory variables, report the labelled \(r\) with sensible rounding, and describe its direction and strength in context. If you report \(r^2\), translate it into a percentage of response-variable variation accounted for by the linear regression. As in “Writing Descriptions in Context,” name the variables and units where appropriate, and avoid causal language unless the study design supports it.

Check Your Understanding

Use the labelled LinReg(a+bx) output and the variable context to answer each question.

  1. What calculator setting must be enabled for a TI-84 to display \(r\) and \(r^2\) with LinReg(a+bx)?
  2. A calculator reports \(b=-1.4\), \(r^2=0.49\), and \(r=-0.7\). Which value is the correlation, and what does its sign indicate?
  3. If \(r^2=0.36\), what percentage of variation in the response is accounted for by the linear regression on the explanatory variable?
  4. Why should the paired explanatory and response values remain on the same rows when entering L1 and L2?
  5. A student says, “\(r^2=0.81\), so 81% of the participants were predicted correctly.” Explain why this is not an appropriate interpretation.