Let the Calculator Report \(r\)
In “Calculating \(r\) From Standardized Values,” you found the correlation by standardizing paired observations and combining their z-scores. For a larger data set, a calculator can do that arithmetic quickly. The TI-84’s LinReg(a+bx) command fits a least-squares regression line and, when diagnostics are on, reports both \(r\) and \(r^2\).
The command’s name describes the line it fits: \(y=a+bx\), where \(a\) is the y-intercept and \(b\) is the slope. These are regression quantities, not correlation. The diagnostic values appear separately. In particular, \(r\) is the correlation between the explanatory and response variables; \(r^2\) is the square of that correlation and is also called the coefficient of determination.
For \(r\) to be defined, both variables must vary. As in “What the Correlation Coefficient Measures,” it summarizes the direction and strength of a linear association between two quantitative variables. Before relying on it, inspect the scatterplot: as discussed in “Matching Correlations to Scatterplots,” a curved pattern or an influential unusual point can make a single correlation misleading.
Turn Diagnostics On and Run LinReg(a+bx)
Some TI-84 calculators have diagnostics turned off initially. In that case, a LinReg command can show \(a\) and \(b\) but omit \(r^2\) and \(r\). Turn diagnostics on once, then run the regression again. The setting generally remains on until it is changed.
Press STAT, choose EDIT, and enter the explanatory-variable values in L1 and their matching response-variable values in L2. Each row must contain a pair from the same individual or case.
Press 2nd, then 0 to open the catalog. Scroll to DiagnosticOn, select it, and press ENTER. Press ENTER again to execute it. The calculator should display “Done.”
Press STAT, open the CALC menu, and select LinReg(a+bx). The menu entry may be farther down the list, so scroll to the command rather than relying on a menu number.
Enter L1, a comma, and L2, then select Calculate or press ENTER. The command-line form is LinReg(a+bx) L1,L2.
Record \(a\), \(b\), \(r^2\), and \(r\). If \(r\) and \(r^2\) are missing, turn on DiagnosticOn and rerun the command.
List order matters. Put the explanatory variable in L1 and the response variable in L2, following the convention from “Making a Scatterplot on the TI-84.” Swapping the lists does not change \(r\) or \(r^2\), but it changes the regression equation: the slope and intercept then describe a different response variable.
Worked Example: A Positive Association
Worked Example: A Positive Association
A fictional class records the number of practice sessions \(x\) and a skills-check score \(y\), in points, for five participants. The values are invented for this example. Enter \(x\) in L1 and \(y\) in L2.
| Participant | Practice sessions \(x\) | Skills-check score \(y\), points |
|---|---|---|
| 1 | 1 | 2 |
| 2 | 2 | 4 |
| 3 | 3 | 3 |
| 4 | 4 | 5 |
| 5 | 5 | 6 |
Calculator procedure. After entering the lists and confirming that DiagnosticOn is enabled, run LinReg(a+bx) L1,L2. The calculator reports \(a=1.3\), \(b=0.9\), \(r^2=0.81\), and \(r=0.9\). The fitted line is \(\hat{y}=1.3+0.9x\), but the correlation is the separately labelled \(r\), not the slope \(b\).
Check the correlation. The means are \(\bar{x}=3\) sessions and \(\bar{y}=4\) points. The sums of squared deviations are \(\sum(x-\bar{x})^2=10\) and \(\sum(y-\bar{y})^2=10\). The sum of paired deviation products is \(9\): the products are \(4,0,0,1,4\). Thus the deviation formula gives:
Squaring gives \(r^2=(0.9)^2=0.81\), which agrees with the calculator. In context, these five participants show a positive linear association between practice sessions and skills-check score: participants with more sessions tended to have higher scores. The fitted line’s slope of \(0.9\) points per session is not the correlation; the correlation \(0.9\) is unitless.
Worked Example: A Negative Association
Worked Example: A Negative Association
A fictional recreation program records the number of days participants used a route-planning app \(x\) and the number of missed turns \(y\) on a practice route. Enter app-use days in L1 and missed turns in L2.
| Participant | App-use days \(x\) | Missed turns \(y\) |
|---|---|---|
| 1 | 1 | 10 |
| 2 | 2 | 8 |
| 3 | 3 | 9 |
| 4 | 4 | 6 |
| 5 | 5 | 7 |
Calculator output. With diagnostics on, LinReg(a+bx) L1,L2 reports \(a=10.4\), \(b=-0.8\), \(r^2=0.64\), and \(r=-0.8\). The slope is negative, which agrees with the downward direction of the association, but its units are missed turns per app-use day. The correlation is unitless.
Check the sign and value. Here \(\bar{x}=3\) days and \(\bar{y}=8\) missed turns. The x-deviations are \(-2,-1,0,1,2\), and the y-deviations are \(2,0,1,-2,-1\). Their paired products sum to \(-8\), while each variable’s squared deviations sum to \(10\). Therefore:
The negative value of \(r\) matches the pattern: participants with more app-use days tended to have fewer missed turns in these observations. The positive value of \(r^2\) does not mean the association is positive. Squaring removes the sign, so use \(r\), together with the scatterplot, to describe direction.
Worked Example: Read Both Diagnostics Correctly
Worked Example: Read Both Diagnostics Correctly
A fictional technology club records weekly practice time \(x\), in hours, and a task score \(y\), in points. A student runs LinReg(a+bx), but the first output shows only \(a\) and \(b\). The student turns on DiagnosticOn as described above and runs the command again.
| Member | Practice time \(x\), hours | Task score \(y\), points |
|---|---|---|
| 1 | 1 | 3 |
| 2 | 2 | 4 |
| 3 | 3 | 5 |
| 4 | 4 | 4 |
| 5 | 5 | 6 |
Read the output. The calculator reports \(a=2.6\), \(b=0.6\), \(r^2\approx0.6923\), and \(r\approx0.8321\), with values rounded to four decimal places where needed. The slope and intercept give the fitted line \(\hat{y}=2.6+0.6x\); \(r\) and \(r^2\) summarize the linear association and its fit.
Verify the diagnostics. The means are \(\bar{x}=3\) hours and \(\bar{y}=4.4\) points. The x-deviations have squared sum \(10\). The y-deviations are \(-1.4,-0.4,0.6,-0.4,1.6\), whose squared deviations sum to \(5.2\). The paired deviation products are \(2.8,0.4,0,-0.4,3.2\), with sum \(6\). So:
Squaring the unrounded value gives \(r^2=36/52=9/13\approx0.6923\). The calculator’s two diagnostic values agree. In context, the five club members show a fairly strong positive linear association between weekly practice time and task score. About \(69.23\%\) of the variation in task scores among these members is accounted for by the least-squares regression of score on practice time. This is a description of the fitted linear relationship, not evidence that practice time caused the scores.
What \(r^2\) Adds
The correlation \(r\) communicates both direction and strength of a linear association. The coefficient of determination \(r^2\) communicates how much of the variation in the response variable is accounted for by its linear regression on the explanatory variable. Because \(r^2\) is a proportion, it ranges from \(0\) to \(1\), and it has no direction or units.
For example, \(r^2=0.64\) means that \(64\%\) of the variation in the response values in the data is accounted for by the least-squares regression line using the explanatory variable. It does not mean that \(64\%\) of individuals are predicted correctly, that the model is accurate for every observation, or that the explanatory variable caused the response. Interpret the percentage in context and keep the response variable clear.
Common Mistakes and AP Exam Tips
- Assuming \(r\) always appears. If diagnostics are off, the calculator may omit \(r\) and \(r^2\). Turn on DiagnosticOn and rerun the regression; do not guess \(r\) from the slope.
- Mixing up \(r\) and \(r^2\). The square \(r^2\) is nonnegative and gives no direction. Use the labelled \(r\) to report whether the linear association is positive or negative.
- Calling the slope a correlation. The slope \(b\) has units, while \(r\) is unitless and always lies between \(-1\) and \(1\). A slope can be much larger than 1 or less than \(-1\).
- Swapping the lists accidentally. Confirm that each L1 entry is paired with the matching L2 entry. Swapping explanatory and response variables changes the regression line, even though \(r\) remains the same.
- Interpreting \(r^2\) as causation or prediction accuracy. State the percent of variation in the response accounted for by the linear regression on the explanatory variable. Do not claim that the relationship is causal or that the same percentage of individual predictions is correct.
- Reporting a correlation without checking the scatterplot. A curved pattern or unusual point can make \(r\) an incomplete summary. Use the scatterplot and DUFS ideas from earlier tutorials to check form and unusual features.
- Rounding too early or misreading the display. Keep the calculator’s full precision for checks, and use the output labels. If you square a rounded display of \(r\), the result may differ slightly from the displayed \(r^2\).
For full-credit communication, identify the response and explanatory variables, report the labelled \(r\) with sensible rounding, and describe its direction and strength in context. If you report \(r^2\), translate it into a percentage of response-variable variation accounted for by the linear regression. As in “Writing Descriptions in Context,” name the variables and units where appropriate, and avoid causal language unless the study design supports it.
Check Your Understanding
Use the labelled LinReg(a+bx) output and the variable context to answer each question.
- What calculator setting must be enabled for a TI-84 to display \(r\) and \(r^2\) with LinReg(a+bx)?
- A calculator reports \(b=-1.4\), \(r^2=0.49\), and \(r=-0.7\). Which value is the correlation, and what does its sign indicate?
- If \(r^2=0.36\), what percentage of variation in the response is accounted for by the linear regression on the explanatory variable?
- Why should the paired explanatory and response values remain on the same rows when entering L1 and L2?
- A student says, “\(r^2=0.81\), so 81% of the participants were predicted correctly.” Explain why this is not an appropriate interpretation.