Tutorials › AP Statistics › Common Communication Errors in Regression

Comparing and communicating regression models · Tutorial 997 of 1000

Common Communication Errors in Regression

Practice replacing common but misleading regression statements with precise, contextual interpretations that match the evidence.

Intermediate 9 min read

What You'll Learn

  • Distinguish what slope, \(r\), \(r^2\), and \(s\) each describe.
  • Add variables, units, cases, and context to incomplete interpretations.
  • Replace causal claims with language appropriate for observational data.
  • Explain why \(r^2\) is not a percentage of predictions that are correct.
  • Communicate an intercept and a fitted value without overstating their meaning.
  • Use side-by-side rewrites to make regression critiques specific and useful.

Small Wording Errors Can Change a Regression Conclusion

A regression response can include accurate output and still give a reader the wrong impression. “The model is 80% accurate,” “the slope is 0.4,” and “each extra hour causes an increase” may sound plausible, but each statement either misinterprets a statistic or leaves out information needed to understand it.

In “Reviewing a Classmate’s Regression Argument,” you practiced identifying what is wrong with a claim and offering a defensible replacement. This tutorial focuses on recurring communication errors and puts weak and improved versions side by side. The goal is not to memorize a perfect sentence for every setting. It is to match each statement to what the regression actually describes, in context and with appropriate limits.

Definition: A regression communication error is wording that misstates a regression result, leaves out context needed to interpret it, or makes a claim that the data and study design do not support.

As in “Turning Output Into a Contextual Interpretation,” keep the roles of the statistics distinct. The slope describes predicted response change per unit of the explanatory variable. The correlation \(r\) describes the direction and strength of a linear association. The coefficient of determination \(r^2\) describes the proportion of variation in the response accounted for by the linear model. The residual standard deviation \(s\) describes the typical size of residuals, in response units.

A Quick Audit Before You Write

Before reporting a result, check the sentence for four things: the variables and units, the statistic’s correct meaning, the cases and setting, and any limit on the claim. Many common errors can be fixed by making one of these elements explicit. For instance, a slope without units is incomplete; a causal verb may overstate an observational association; and a claim about “all” cases may go beyond the data.

1
Name what was measured.
Identify the explanatory variable \(x\), response variable \(y\), their units, and the cases represented.
2
Use the statistic for its proper job.
Do not substitute \(r\) for slope, interpret \(r^2\) as a success rate, or describe \(s\) as a guaranteed prediction error.
3
Match the strength of the claim to the design.
Describe an association for observational data; do not claim cause and effect unless the study design supports it.
4
Check the model’s scope.
Keep conclusions tied to the cases, setting, and observed explanatory-variable range represented in the analysis.

Worked Rewrites: Slope, \(r^2\), and \(s\)

The first example combines several errors that often appear together. Read each weak statement separately: fixing one does not automatically fix the others.

Worked Example: Practice Time and Free-Throw Percentage

Situation. In an invented observational project, students recorded weekly basketball practice time and free-throw percentage for 40 players in one club. Let \(x\) be practice time in minutes per week and \(y\) be free-throw percentage. Practice time was chosen by the players, not assigned by researchers. The observed practice times ranged from 20 to 180 minutes. A regression output summary gives

$$ \hat{y}=41+0.35x,\qquad r=0.72,\qquad r^2=0.5184,\qquad s=5.1\text{ percentage points}. $$

A report says: “Every extra minute of practice causes a 0.35% improvement. The model predicts 52% of players correctly, and its errors are always about 5.1 percentage points.”

State. The report makes an unsupported causal claim, gives an incomplete interpretation of the slope, misinterprets \(r^2\), and treats \(s\) as a guaranteed error size.

Plan. Interpret the slope with the response and explanatory-variable units. Convert \(r^2\) to a percentage and state which variation it describes. Then explain what \(s\) says about residuals and why it does not guarantee a particular prediction’s accuracy.

Do. The slope is \(0.35\) free-throw percentage points per additional minute of weekly practice. A suitable interpretation is: “Among these 40 players, each additional minute of weekly practice was associated with a predicted increase of 0.35 percentage points in free-throw percentage, according to the fitted line.” The phrase “causes an improvement” is not supported: the players chose their practice time, and other differences between players could be related to both practice and free-throw percentage.

To express \(r^2\) as a percentage, calculate \(0.5184\times100\%=51.84\%\), or about \(51.8\%\) rounded to one decimal place. This means that about 51.8% of the variation in free-throw percentages among these players is accounted for by the linear regression of free-throw percentage on weekly practice time. It does not mean that the model predicts 52% of players correctly.

The reported \(s=5.1\) percentage points describes the typical size of residuals in this setting. Equivalently, observed free-throw percentages typically differ from their fitted values by about 5.1 percentage points. It does not mean that every prediction is wrong by exactly 5.1 points or that all predictions are within 5.1 points of the observed percentages.

Conclude. A more careful summary is: “For the 40 players in this club, weekly practice time and free-throw percentage had a positive linear association. The fitted line predicts an increase of 0.35 percentage points in free-throw percentage per additional minute of weekly practice. About 51.8% of the variation in free-throw percentage among these players is accounted for by the linear model, and the typical residual size is 5.1 percentage points. Because practice time was not assigned, these data do not establish that more practice causes a higher free-throw percentage.”

Notice how the revised version does more than replace “causes” with “is associated with.” It also supplies the units, identifies which response varies, and specifies the players represented. Those details make the statistical meaning clear rather than leaving the reader to guess.

Do Not Confuse Slope, Correlation, and Percentage Points

A slope and a correlation may both be positive or both be negative, but they are not interchangeable. Slope has units: response units per explanatory-variable unit. The correlation \(r\) has no units and is bounded between \(-1\) and \(1\). Calling \(r\) a “change per unit” gives it the meaning of a slope when it does not have one.

Another wording trap occurs when a response is recorded as a percentage. A change from 70% to 71% is an increase of one percentage point. It is not a 1% increase relative to the original value. When the slope predicts a change in a percentage-valued response, percentage points are often the clearest units.

Worked Example: Distance and E-Bike Battery Charge

Situation. In an invented analysis, 31 e-bike rides were recorded. Let \(x\) be kilometers traveled since the battery was fully charged and \(y\) be the remaining battery charge, recorded as a percentage. The observed distances ranged from 4 to 48 kilometers. The fitted line is \(\hat{y}=96-1.4x\), with \(r=-0.90\), \(r^2=0.81\), and \(s=6.2\) percentage points.

A student writes: “The correlation means the battery loses 0.90% per kilometer. At zero kilometers, riders usually have 96% battery. After 30 kilometers, every bike will have 54% charge, with a 6.2% error.”

Plan. Separate correlation from slope, check whether zero kilometers is within the observed range before interpreting the intercept practically, calculate the fitted value at 30 kilometers, and describe \(s\) without implying a guarantee.

Do. The correlation \(r=-0.90\) indicates a strong negative linear association between distance traveled and remaining battery charge for these rides. The slope, not \(r\), is \(-1.4\) percentage points per kilometer. In context, each additional kilometer traveled is associated with a predicted decrease of 1.4 percentage points in remaining charge according to the fitted line.

At zero kilometers, the equation gives \(\hat{y}=96-1.4(0)=96\%\). This is the model’s predicted charge at zero kilometers, but zero is below the observed range of 4 to 48 kilometers. Thus, treating 96% as an established typical starting charge for these rides goes beyond the observed distances.

For a 30-kilometer ride, the fitted value is

$$ \hat{y}=96-1.4(30)=96-42=54\%. $$

Thirty kilometers is within the observed distance range, so this is not extrapolation. It is still a fitted value, not a guarantee that every bike traveling 30 kilometers will have exactly 54% charge remaining. The statistic \(s=6.2\) percentage points means that residuals typically have a size of about 6.2 percentage points in this data set. It does not specify the error for this particular ride.

Conclude. A defensible rewrite is: “For the 31 recorded rides, distance traveled had a strong negative linear association with remaining battery charge. The fitted line predicts a decrease of 1.4 percentage points in remaining charge per additional kilometer. At 30 kilometers, the predicted remaining charge is 54%; the typical residual size is 6.2 percentage points. The intercept is the fitted value at zero kilometers, which is below the observed distance range, so it should not be presented as a typical observed charge.”

Use Each Statistic to Support the Right Claim

A number can be accurate but still be used to support the wrong statement. The earlier tutorial “Selecting Evidence for a Written Conclusion” emphasized matching a claim to evidence. For instance, \(r\) addresses the direction and strength of a linear association; it does not give the predicted response change. A residual plot can show a pattern left by a fitted line; it does not establish that an explanatory variable causes the response.

The table below shows common wording problems and the kind of repair that makes the communication more accurate. The improved statements still need context-specific details, such as the cases and setting, when those details are available.

Weak or incorrect wordingImproved communication
“\(r=-0.7\), so \(y\) drops 0.7 units each time \(x\) rises by one.”“The correlation indicates a negative linear association. Use the slope, with its units, to describe predicted response change per unit of \(x\).”
“\(r^2=0.60\), so the model is correct 60% of the time.”“About 60% of the variation in the response among the cases used to fit the model is accounted for by the linear regression on \(x\).”
“\(s=3\), so every prediction is within 3 units.”“Residuals typically have a size of about 3 response units; this does not guarantee the error for an individual prediction.”
“The slope proves that changing \(x\) will change \(y\).”For observational data, describe the predicted association and avoid a causal conclusion unless the study design supports one.
“The intercept is the usual response.”State that it is the model’s predicted response at \(x=0\), then check whether zero is in the observed range and meaningful in context.

Worked Example: Tree Cover and Afternoon Temperature

Situation. In an invented observational analysis, 35 city blocks in one district were studied during one summer. Let \(x\) be the percentage of a block covered by tree canopy and \(y\) be afternoon temperature in degrees Celsius. Canopy cover ranged from 8 to 62 percentage points. The fitted line is \(\hat{y}=31.6-0.08x\), with \(r=-0.74\), \(r^2=0.5476\), and \(s=1.9^\circ\text{C}\).

A draft report says: “The correlation shows that each extra percentage point of tree cover lowers temperature by 74%. Since \(r^2=-0.55\), trees explain 55% of the temperature drop. This proves adding trees will lower temperatures in every city. The model is off by 1.9 degrees.”

State. The draft confuses \(r\) with slope, gives \(r^2\) the wrong sign and an unclear meaning, makes a causal claim from observational data, generalizes beyond the studied setting, and describes \(s\) imprecisely.

Plan. Interpret the slope using degrees Celsius per percentage point of canopy cover. Calculate and interpret \(r^2\) as a proportion of response variation. Describe \(s\) as typical residual size, then limit the conclusion to the blocks and setting represented.

Do. The slope is \(-0.08^\circ\text{C}\) per percentage point of tree canopy. It predicts that each additional percentage point of canopy cover is associated with a decrease of \(0.08^\circ\text{C}\) in afternoon temperature, according to the fitted line. The value \(r=-0.74\) describes the direction and strength of the linear association; it is not the predicted temperature change per percentage point.

The calculation \(r^2=(-0.74)^2=0.5476\) confirms that \(r^2\) is positive. As a percentage, \(0.5476\times100\%=54.76\%\), or about \(54.8\%\) rounded to one decimal place. Thus, about 54.8% of the variation in afternoon temperatures among these 35 blocks is accounted for by the linear regression of temperature on tree-canopy cover. It does not mean trees explain 54.8% of each temperature drop.

The \(s\) value means residuals typically have a size of about \(1.9^\circ\text{C}\). Equivalently, observed afternoon temperatures typically differ from their fitted values by about \(1.9^\circ\text{C}\). That is a description of typical scatter, not a guarantee about every block. Because the data are observational and come from one district during one summer, they do not prove that increasing tree cover causes temperatures to fall or establish that the same relationship applies in every city.

Conclude. A stronger version is: “Among the 35 blocks studied in this district during one summer, tree-canopy cover and afternoon temperature had a negative linear association. The fitted line predicts a decrease of \(0.08^\circ\text{C}\) in afternoon temperature per additional percentage point of canopy cover. About 54.8% of the variation in afternoon temperature among these blocks is accounted for by the linear model, and the typical residual size is \(1.9^\circ\text{C}\). These observational data do not establish that adding trees causes a temperature decrease, and the results should not automatically be generalized to other cities.”

Common Mistakes and AP Exam Tips

  • Copying a statistic without its meaning. “The slope is \(-0.08\)” is not a contextual interpretation. State the predicted response change, the direction, the explanatory-variable unit, and the response unit.
  • Calling \(r^2\) an accuracy rate. A regression model does not classify cases as correct or incorrect in the way that wording suggests. Identify the response variation and the linear regression that accounts for the stated proportion.
  • Using correlation as a rate of change. \(r\) has no units. Use the slope to describe predicted change per unit of \(x\); use \(r\) to describe direction and strength of a linear association.
  • Calling association causation. A regression slope is not automatically a causal effect. For observational data, write “is associated with” or “the line predicts,” and explain that the data do not establish cause and effect.
  • Turning \(s\) into a promise. Do not say every observed value is within \(s\) of its fitted value. Say that residuals typically have a size of about \(s\) response units, or that observed responses typically differ from fitted responses by about \(s\) response units.
  • Giving an intercept a practical meaning without checking \(x=0\). The intercept is the predicted response at zero explanatory-variable units. Check whether zero is in the observed range and whether that value makes sense in the situation.
  • Writing a conclusion that is broader than the data. A result from one group or setting does not automatically apply elsewhere. State which cases and setting the analysis represents, and check the observed \(x\)-range when discussing a prediction.

A strong AP response is not necessarily long. It is specific. Name the statistic, say what it describes, include the relevant context and units, and avoid claims stronger than the design or observed data support. When you critique a weak statement, identify the exact error and offer a replacement sentence rather than merely writing “incorrect.”

Key takeaway: Regression communication improves when every statistic is used for its proper purpose and each claim includes context, units, and appropriate limits. A precise rewrite explains what the model supports without turning association into causation or typical scatter into a guarantee.

Check Your Understanding

For each statement, identify the communication error and describe what a more complete version should say.

  1. A regression of weekly reading time on quiz score has slope \(1.6\). What context and units are needed to interpret this slope?
  2. A report gives \(r=-0.83\) and says, “Quiz scores fall by 0.83 points for each extra hour of reading.” Which statistic should be used for predicted change per hour, and what does \(r\) describe?
  3. A model has \(r^2=0.42\). A student says, “It predicts 42% of students correctly.” What does \(r^2\) describe instead?
  4. In observational data, a positive association is found between daily water use and garden size. Why should a report be cautious about saying that increasing garden size causes water use to rise?
  5. A regression output reports \(s=2.4\) units. Write a careful statement about what this value says and what it does not guarantee.