Finding the Right Number in a Regression Summary
In “\(r^2\) Is Not Percent Correct Predictions,” you learned what the coefficient of determination describes. Now the task is to locate that value in computer output and avoid selecting a nearby statistic with a similar name. Software may display several summaries together, including \(R\)-sq, adjusted \(R\)-sq, and a correlation measure. They are related, but they are not interchangeable.
For simple linear regression, the coefficient of determination is written \(r^2\) in many statistics lessons and may be labeled R-squared, R Square, or R-sq in output. The capital \(R\) in the software label does not change what the statistic means in this setting. Look for the label itself rather than choosing a number just because it is near the regression coefficients.
A coefficient of determination of \(0.72\) and one reported as \(72\%\) represent the same amount. Before interpreting the result, check whether the output uses a decimal or percentage scale. Then name the response and explanatory variable in your interpretation, as in “Interpreting \(r^2\) in Context.”
Keep Three Output Entries Separate
A regression summary can show \(R\)-sq beside adjusted \(R\)-sq and a correlation measure. The coefficient of determination for the standard AP Statistics interpretation is the value labeled \(R\)-sq, not adjusted \(R\)-sq. Adjusted \(R\)-sq is a modified summary that accounts for the number of predictors and the amount of data. It can be useful in other model-comparison settings, but it is not the usual \(r^2\) to report for interpreting a simple linear regression in this course.
Correlation, \(r\), is a different statistic. It describes the direction and strength of a linear relationship and can be positive or negative. The coefficient of determination is nonnegative. In simple linear regression, \(r^2\) is the square of the correlation, but that relationship does not make a printed correlation and a printed \(R\)-sq the same number or the same statistic. A report of \(r=-0.76\), for example, includes direction; an \(R\)-sq value does not.
Labels vary by program. Some software prints the signed Pearson correlation as Correlation or \(r\); other output may show a nonnegative field called Multiple R. Read the label and the software’s convention. A number labeled “Multiple R” should not automatically be treated as the signed correlation. If a signed correlation is needed, look for an explicitly signed correlation result or consult the relevant context and output conventions.
Worked Example: Locate R-sq in a Summary Table
A fictional school project uses the number of minutes students spend practicing a keyboarding lesson to predict their typing speed, in words per minute. Software gives the following summary for 51 students.
| Output entry | Value |
|---|---|
| Multiple R | 0.9100 |
| R Square | 0.8281 |
| Adjusted R Square | 0.8246 |
| Observations | 51 |
State. Identify the coefficient of determination and interpret it in context.
Plan. Select the entry labeled “R Square,” not “Adjusted R Square” or “Multiple R.” The response is typing speed, and the explanatory variable is practice time. Convert the decimal to a percentage for the interpretation.
Do. The output reports \(R\)-sq \(=0.8281\). As a percentage, \(0.8281(100\%)=82.81\%\), or about \(82.8\%\). The adjusted value \(0.8246\) is a separate, modified summary. “Multiple R” is also a separate output entry.
Conclude. About 82.8% of the variation in students’ typing speeds is accounted for by the linear relationship between typing speed and minutes spent practicing. This does not mean that 82.8% of students had correct predictions, and it is not the adjusted \(R\)-sq value.
The entries may be close, as they are here, but closeness does not make them interchangeable. Answering with adjusted \(R\)-sq would give a different value from the requested coefficient of determination, even if both values are useful in some software or modeling contexts.
Read the Label Before Interpreting the Number
One useful habit is to scan the output in two passes. First, identify the statistic by its label. Second, identify the scale and the variables in the model. This prevents a decimal from being mistaken for a percentage and prevents a value beside \(R\)-sq from being assigned the wrong meaning.
Locate \(R\)-sq, \(R\) Square, or \(R\)-squared. Keep adjusted \(R\)-sq and correlation entries separate.
Decide whether the output gives a decimal such as \(0.58\) or a percentage such as \(58\%\).
Identify which variable is the response and which is the explanatory variable before writing the interpretation.
Describe variation in the response accounted for by its linear relationship with the explanatory variable—not a percentage of cases or correct predictions.
Worked Example: Distinguish R-sq From Correlation
A fictional environmental club examines the relationship between weekly rainfall and the amount of water collected by a small rain barrel, measured in liters. A statistics program reports \(r=-0.760\), \(R\)-squared \(=0.5776\), and adjusted \(R\)-squared \(=0.5600\). Which number is the coefficient of determination, and what does it mean?
State. The coefficient of determination is the value labeled \(R\)-squared, \(0.5776\). The reported correlation, \(r=-0.760\), and adjusted \(R\)-squared, \(0.5600\), are different entries.
Plan. Interpret \(0.5776\) as a percentage of variation in the response, water collected. Include its linear relationship with the explanatory variable, weekly rainfall. Do not use the negative sign from \(r\) as if it were part of \(R\)-squared.
Do. Convert the coefficient to a percentage: \(0.5776(100\%)=57.76\%\), or about \(57.8\%\). The correlation’s negative sign describes direction; \(R\)-squared is nonnegative. The adjusted value is not the requested coefficient of determination.
Conclude. About 57.8% of the variation in the amount of water collected is accounted for by its linear relationship with weekly rainfall. The reported correlation of \(-0.760\) describes a negative direction, while \(R\)-squared gives the fraction of response variation accounted for by the linear relationship.
Be cautious about a printed value labeled Multiple R. In some software, that entry is nonnegative, even when the relationship is negative. It is not automatically the signed correlation \(r\). The correlation and the coefficient of determination answer different questions, so preserve each value’s label and sign when describing them.
When Output Shows a Percentage
Some programs report \(R\)-sq as a decimal, while others display a percent or round the value to a few digits. The interpretation should match the reported precision. For instance, \(68.0\%\) is already on the percentage scale; do not multiply it by 100 again. If a value is shown as \(0.680\), convert it to \(68.0\%\).
Rounding can make two related entries look less consistent than they were before rounding. For example, software might show correlation to three decimal places and \(R\)-sq to three decimal places. Do not use rounded values to challenge a more precise output value or to claim that two entries must match exactly. For the task “find \(R\)-sq,” report the labeled value at the precision shown and interpret it appropriately.
Worked Example: Read a Rounded Output Carefully
A fictional transit study uses trip distance to predict the time a shuttle takes to complete a route. A program summary displays “R-sq = 68.0%,” “Adj R-sq = 66.9%,” and “Multiple R = 0.825.” The fitted slope for distance is negative. Identify the coefficient of determination and write an interpretation without confusing it with the other entries.
State. The coefficient of determination is the output value labeled “R-sq,” which is \(68.0\%\). The adjusted value and “Multiple R” are not the requested entry.
Plan. Treat the displayed value as a percentage, not a decimal. Name route time as the response and trip distance as the explanatory variable. Do not call the nonnegative “Multiple R” entry the signed correlation.
Do. The output already reports \(R\)-sq as \(68.0\%\), so no multiplication by 100 is needed. The adjusted \(R\)-sq is \(66.9\%\), which is a different summary. “Multiple R” is printed as \(0.825\), without a negative sign; the slope’s negative sign warns against reading that field as a signed correlation.
Conclude. About 68.0% of the variation in shuttle route times is accounted for by the linear relationship between route time and trip distance. The \(66.9\%\) adjusted value is not the \(R\)-sq requested, and the \(0.825\) “Multiple R” entry should not be reported as a signed correlation.
Common Mistakes and AP Exam Tips
- Choosing adjusted \(R\)-sq. It is often printed immediately beside \(R\)-sq, but it is a different, adjusted summary. For the standard simple-regression coefficient of determination, read the entry labeled \(R\)-sq or \(R\) Square.
- Calling \(R\)-sq the correlation. Correlation \(r\) can be negative or positive; \(R\)-sq is nonnegative. Use the statistic’s label, and preserve the sign when the output explicitly reports a signed correlation.
- Misreading a percentage as a decimal. If \(R\)-sq is \(0.68\), that is \(68\%\). If it is already reported as \(68\%\), do not multiply again.
- Giving a context-free interpretation. A full-credit interpretation names the response and explanatory variable and says that the percentage is variation in the response accounted for by their linear relationship.
- Describing cases instead of variation. Avoid phrases such as “68% of the trips were predicted correctly.” \(R\)-sq does not count observations or exact predictions.
- Assuming all software labels mean the same thing. Read the exact label, especially for fields such as “Multiple R.” Do not assume that an unsigned output entry is the signed correlation.
On an AP response, show that you found the requested statistic by naming its output label, then interpret its value in context. If the question asks for \(r^2\), do not report adjusted \(R\)-sq or correlation in its place. If it asks for correlation, do not report \(R\)-sq as though it included direction.
Check Your Understanding
Use the labels and values in each item to identify the coefficient of determination and describe what it means.
- A model predicts plant height from the number of hours of light. Output lists “R Square = 0.64” and “Adjusted R Square = 0.61.” Which value is the coefficient of determination? Write its interpretation in context.
- A program reports a correlation of \(-0.70\) and \(R\)-squared of \(0.49\). Which entry has a direction sign? What does the \(R\)-squared value describe?
- An output shows “R-sq = 72%.” Should you multiply this value by 100 before interpreting it? Explain.
- A summary lists “Multiple R = 0.83” and a negative fitted slope. Why should you avoid automatically calling \(0.83\) the signed correlation?
- A student says, “An \(R\)-sq of \(0.58\) means 58% of the observations were predicted correctly.” Identify the error and write the kind of statement that \(R\)-sq supports.