Tutorials › AP Statistics › Explaining Regression to a Nontechnical Audience

Comparing and communicating regression models · Tutorial 992 of 1000

Explaining Regression to a Nontechnical Audience

Learn a practical way to explain a sleep study’s regression findings in plain language without losing their meaning or overstating what they show.

Intermediate 9 min read

What You'll Learn

  • Translate a regression slope into a sentence with context and units.
  • Explain what the correlation and the proportion of variation accounted for say about the study.
  • Describe typical prediction error without implying an exact guarantee.
  • Distinguish an observed association from a cause-and-effect claim.
  • Adapt a technical regression summary for a general reader.

Make the Meaning Clear, Not Just the Vocabulary

Regression results often arrive as an equation, a correlation, and measures of fit. A general reader may not know those terms, but replacing them with vague phrases can make the results less accurate. The goal is to translate the statistical meaning into familiar language while keeping the context, units, and limits of the study.

In “Describing Limitations of a Regression Analysis,” you learned to connect a conclusion to the data and the study’s scope. That matters just as much when writing for a nontechnical audience: clear wording should not turn an association into a cause, or make a prediction sound more certain than it is.

Throughout this tutorial, consider an invented observational study of 80 volunteer high-school students from two schools. Researchers recorded each student’s sleep duration for one night and a score on a next-morning attention task scored from 0 to 100. Students in the study slept between 5 and 9 hours. The reported regression results are:

$$ \widehat{\text{attention score}}=48+4.2(\text{hours of sleep}),\qquad r=0.70,\qquad r^2=0.49,\qquad s=8\text{ points}. $$

These results are invented for practice. The equation and statistics summarize the study’s sample; they do not, on their own, establish what would happen if a student deliberately slept more.

Definition: Plain-language regression communication explains what a model’s results mean in familiar words while retaining the correct variables, units, scope, and uncertainty.

A Translation Ladder for Regression Results

A useful translation has four parts: identify the statistical claim, keep the quantity and its units, replace technical wording with everyday language, and add any limit that changes how a reader should understand the claim. You do not need to include every statistic in every short report. Choose the ones that answer the reader’s question.

1
Name what is being compared.
Say which measured feature is being used to describe or predict which outcome. In this study, the comparison is between hours of sleep and attention-task score the next morning.
2
Translate the number without changing its meaning.
Give the slope as a predicted score difference for a one-hour difference in sleep. Keep the units: hours of sleep and points on the task.
3
Explain how well the line summarizes the data.
Describe \(r\), \(r^2\), or \(s\) only if they help answer the question. Each describes a different feature of the pattern or its scatter.
4
Keep the conclusion within the study’s reach.
Identify the participants or observed range when relevant, and do not write that sleep caused a higher score when the study only observed an association.

The regression line predicts an attention score from a sleep duration. Its slope, \(4.2\), means that for each additional hour of sleep, the model predicts an attention score 4.2 points higher. Because this study is observational, this is a description of the association in the data, not a claim that gaining an hour of sleep causes a score increase.

The correlation \(r=0.70\) summarizes the direction and strength of the linear association: students with longer recorded sleep tended to have higher attention scores. The value \(r^2=0.49\) says that 49% of the variation in attention scores among these students is accounted for by the linear model using sleep duration. It does not mean that sleep caused 49% of each student’s score, or that the model predicts every student’s score correctly.

The residual standard deviation \(s=8\) points describes the typical vertical distance between observed scores and the scores predicted by the line. In everyday wording, observed attention scores typically differed from the line’s predictions by about 8 points. This is a summary of the residual scatter, not a promise that each prediction will be within 8 points.

Use Familiar Words Without Hiding the Statistics

Some technical terms are useful if they are translated at the moment they appear. Other terms can often be omitted from a brief public summary. The table shows ways to preserve meaning without making readers decode a list of statistics.

Technical wordingPlain-language wordingMeaning to preserve
Positive linear associationStudents with longer sleep tended to have higher scoresDescribe a pattern, not a cause
Slope of 4.2 points per hourThe line predicts 4.2 more score points for each additional hour of sleepKeep the direction and both units
\(r^2=0.49\)The line accounts for 49% of the differences in scores in this sampleRefer to variation among scores, not a share of an individual score
Residual standard deviation \(s=8\)Scores typically differed from the line’s predictions by about 8 pointsDescribe typical scatter, not a guaranteed error limit

“Plain language” does not mean removing every number. A general reader can understand a number when the sentence explains what it measures. “The slope was 4.2” is difficult to interpret; “the line predicts 4.2 more attention-score points for each additional hour of sleep” supplies the outcome, direction, and units.

Worked Example: Rewrite a Technical Summary

Worked Example: Sleep Duration and Attention Scores

Original AP-style question. Using the invented study described above, rewrite this technical summary for a general reader: “There was a positive linear association between sleep duration and attention score (\(r=0.70\)). The least-squares regression line had a slope of 4.2 points per hour, \(r^2=0.49\), and residual standard deviation \(s=8\) points.” Include a relevant limitation.

Solution. First identify what each result contributes. The positive value of \(r\) indicates that longer sleep tended to occur with higher scores. The slope gives the model’s predicted score difference per additional hour. The value of \(r^2\) summarizes how much score variation the line accounts for, while \(s\) gives the typical size of the residuals in points.

A plain-language version is: “Among 80 volunteer students from two schools, those who slept longer on the recorded night tended to score higher on the attention task the next morning. The fitted line predicts about 4.2 more points for each additional hour of sleep. Sleep duration accounted for 49% of the variation in scores in this group, and actual scores typically differed from the line’s predictions by about 8 points. Because the study observed students rather than assigning their sleep, it does not show that sleeping longer caused higher attention scores.”

This version translates the statistics but keeps their meanings. It names the participants, gives the slope’s units, explains what the 49% refers to, and treats the 8-point scatter as typical rather than guaranteed. The final sentence limits the conclusion to an association, as appropriate for observational data.

Worked Example: Explain a Difference Between Two Sleep Durations

Worked Example: Compare the Line’s Predictions

Original AP-style question. Use the fitted line to compare the predicted attention scores for students who slept 6 hours and 8 hours. Explain the result for a general reader and avoid implying that sleep caused the difference.

Plan. Substitute each sleep duration into the regression equation. Then compare the two fitted values. Describe the result as a difference between model predictions, not a guaranteed difference between two particular students.

Do. For 6 hours, the predicted score is

$$ \hat{y}=48+4.2(6)=48+25.2=73.2\text{ points}. $$

For 8 hours, the predicted score is

$$ \hat{y}=48+4.2(8)=48+33.6=81.6\text{ points}. $$

The difference between these fitted values is \(81.6-73.2=8.4\) points. This also agrees with the slope: a 2-hour difference corresponds to a predicted difference of \(4.2(2)=8.4\) points.

Conclude. The line predicts an attention score 8.4 points higher for 8 hours of sleep than for 6 hours. Both sleep durations are within the study’s observed range of 5 to 9 hours, but the result is still a model-based comparison, not a guarantee for any two individuals. Since sleep was observed rather than assigned, the study does not show that the extra sleep caused an 8.4-point increase.

Worked Example: Choose What to Include in a Short Report

Worked Example: Write a Brief Public Summary

Original AP-style question. A school newsletter wants a two-sentence summary of the invented sleep study. The editor asks for the main finding and one caution, but does not want a string of unexplained statistics. Draft an accurate summary.

Solution. A short report should include the direction and size of the association, explain the population represented, and state the key caution. The slope is useful because it gives the size of the model’s predicted change in score for an additional hour of sleep. The sample description and observational design matter because they restrict how readers should interpret or apply the result.

A suitable summary is: “In this study of 80 volunteer students from two schools, longer sleep on the recorded night was associated with higher attention-task scores the next morning; the fitted line predicted about 4.2 more points per additional hour. The study describes these students and does not show that more sleep caused higher scores, so the result should not be treated as a guaranteed benefit for every student.”

This summary leaves out \(r\), \(r^2\), and \(s\), not because they are unimportant, but because the editor requested a very brief account. If a reader needs to judge how closely the line fits the scores, \(r^2\) and \(s\) could be added with explanations. A useful report selects statistics for the reader’s question instead of listing every available result.

Common Mistakes and AP Exam Tips

  • Using “increases” as if the study proved a cause. In an observational study, say “was associated with,” “tended to have,” or “the line predicts.” A causal statement needs support beyond a fitted regression line.
  • Giving the slope without its units. “The slope is 4.2” does not say what changes. A complete interpretation names both quantities: the predicted attention score is 4.2 points higher per additional hour of sleep.
  • Calling \(r^2\) the percentage caused by sleep. It is the proportion of variation in the response accounted for by the linear model in the data. It is not a percentage of an individual score or proof of causation.
  • Turning \(s\) into a guarantee. “Predictions are off by exactly 8 points” is too strong. Say that observed scores typically differed from the line’s predicted scores by about 8 points.
  • Reporting the intercept as if it were automatically meaningful. The equation predicts 48 points at zero hours of sleep, but this study observed sleep durations only from 5 to 9 hours. Do not give the intercept a real-world interpretation for zero hours without considering that range and context.
  • Removing the study context to save words. “Sleep improves attention” is shorter, but it overstates the result and leaves out who was studied. Keep the sample and the observational nature of the study when they affect the claim.

For full credit on an AP response, connect the number to the variables and situation. State the slope with direction and units; explain \(r^2\) as variation in the response accounted for by the model; describe \(s\) as typical residual scatter in response units; and qualify an observational association rather than claiming cause and effect. For a general reader, the same accuracy applies—just use familiar words and explain the numbers where they appear.

Key takeaway: A strong plain-language regression summary preserves the result’s context, direction, units, and limits. Translate jargon into meaning, not into a stronger claim.

Check Your Understanding

Use the invented sleep-study results in this tutorial to answer each question in clear, complete language.

  1. Write a plain-language interpretation of the slope \(4.2\) that includes the response units.
  2. What does \(r^2=0.49\) describe in this study? Name one interpretation it does not support.
  3. Explain \(s=8\) points without making it sound like every prediction is within 8 points.
  4. Write one sentence that reports the sleep-attention association without claiming that sleep caused higher scores.
  5. Why should the fitted value at zero hours of sleep be treated cautiously in this study?