Tutorials › AP Statistics › Individual Predictions Versus Mean Predictions

Extrapolation and prediction limits · Tutorial 955 of 1000

Individual Predictions Versus Mean Predictions

Understand what a regression line predicts for one case versus a group average, and why averaging independent responses can make a prediction more precise.

Intermediate 9 min read

What You'll Learn

  • Distinguish a prediction for one individual from a prediction of the mean response at a given explanatory-variable value.
  • Explain why individual responses can vary around the fitted line even when the model predicts their average well.
  • Use residual standard deviation to describe the typical scatter for an individual prediction.
  • Compare the approximate scatter of one response with the scatter of an average of several independent responses.
  • Identify conditions that make averaging more or less effective.
  • Avoid applying a group-level model’s residual standard deviation to predictions about individuals.

One Prediction, Two Different Questions

In “Using Residual Standard Deviation to Gauge Prediction Error,” we used \(s\) to describe the typical size of observed responses’ residuals around a fitted regression line. The next question is: what are we predicting? A model can be used to predict a response for one individual, or to estimate the average response for a group of individuals with a particular explanatory-variable value. Those are different goals, and the average is generally more predictable than any one individual.

For a given value of \(x\), the fitted value \(\hat{y}\) is the line’s predicted response. It estimates the mean response among cases with that \(x\), under the linear model. It can also serve as a point prediction for one individual with that \(x\). But a point prediction does not mean every individual response equals \(\hat{y}\). Individual responses vary around the line.

Definition: An individual prediction uses a fitted regression line to predict the response for one case at a specified \(x\)-value. A mean prediction uses the line to estimate the average response for a group of cases at that \(x\)-value. The fitted value may be the same in both cases, but the amount of response variation around it is not.

Think of the fitted line as giving the center of the response pattern at each \(x\). A single person, object, or event can land noticeably above or below that center. If we average responses from several comparable individuals, some deviations above the line and some below it can offset one another. That tends to make the average closer to the line than an individual response is.

This distinction is about precision, not about changing the line’s prediction. If the line predicts 42 centimeters at a certain \(x\), that is the model’s predicted center for that value. It does not predict that every individual will be 42 centimeters tall. The line’s predicted average may be useful even when the prediction for any particular individual is fairly uncertain.

Why Averages Tend to Vary Less

The residual standard deviation \(s\), as discussed in the previous tutorial, gives a response-unit scale for the typical scatter of individual observed responses around the line. When predicting one individual, that scatter is relevant: the individual’s response can differ from the fitted value by an amount that may be small or large relative to \(s\).

Now suppose we average \(m\) independent responses for individuals at the same \(x\)-value, or at a sufficiently similar value that they share the same model prediction. If the individuals have roughly the same response scatter, positive and negative residuals tend to cancel in the average. The approximate standard deviation of the average’s residuals is \(s/\sqrt{m}\). This is a scale for the average’s typical variation around the line, not a guaranteed error bound.

$$ \text{Approximate scatter for an average of }m\text{ independent responses} \approx \frac{s}{\sqrt{m}} $$

For example, averaging 9 comparable independent responses reduces this approximate scatter to \(s/3\), not to zero. Averaging 25 reduces it to \(s/5\). The improvement gets smaller as the group grows: increasing the group from 9 to 16 changes the multiplier from \(1/3\) to \(1/4\), rather than cutting the scatter in half.

This comparison relies on assumptions. The responses should be independent, the individuals should be comparable in the relevant way, and their scatter around the line should be reasonably similar. If observations are strongly related—for example, repeated measurements on one person or students from the same closely connected household—the deviations may not cancel as much as the calculation suggests.

Conditions for the averaging comparison: The \(s/\sqrt{m}\) rule is an approximate way to compare residual scatter when averaging \(m\) independent, comparable responses. It assumes a similar amount of scatter for the individuals being averaged. It is not a prediction interval, does not account for every source of uncertainty in the fitted line, and does not guarantee how close a particular average will be.

The rule is useful for understanding why a mean prediction can be more precise, but do not turn it into a percentage or a guarantee. In particular, saying that an average’s typical scatter is \(s/\sqrt{m}\) does not mean every group average is within that distance of the line.

Worked Examples: Individual Responses and Averages

Worked Example: One Plant Versus an Average of Nine

Hypothetical setting. A greenhouse team models plant height \(y\), in centimeters, using a watering measure \(x\). At a particular watering level, the fitted line predicts a height of 42 centimeters. The residual standard deviation is \(s=6\) centimeters. Compare a prediction for one plant with a prediction for the average height of 9 independent, comparable plants at that watering level.

State. The fitted value of 42 centimeters is the predicted center of the heights at this watering level. A single plant can vary around that center; the average of 9 plants should vary less if the plants’ responses are independent and have similar scatter.

Plan. Use \(s\) as the typical residual scale for an individual response. For the average of 9 independent responses, calculate the approximate scale \(s/\sqrt{9}\). This comparison describes typical scatter, not a guaranteed range for a plant or an average.

Do. For one plant, the line’s prediction is 42 centimeters, and the typical residual scale is about 6 centimeters. For an average of 9 plants:

$$ \frac{s}{\sqrt{m}} =\frac{6}{\sqrt{9}} =\frac{6}{3} =2\text{ centimeters} $$

The model’s predicted average height is still 42 centimeters. The approximate residual scatter for the average is 2 centimeters, compared with 6 centimeters for one response.

Conclude in context. At this watering level, the fitted line predicts a height of 42 centimeters. A single plant’s height typically varies around the line on a scale of about 6 centimeters. The average height of 9 independent, comparable plants has a smaller approximate scatter scale of 2 centimeters. This does not guarantee that every plant is within 6 centimeters, or that the nine-plant average is within 2 centimeters of 42.

Worked Example: Predicting One Teen’s Sleep Versus a Group Average

Hypothetical setting. A researcher records evening screen time \(x\), in hours, and sleep duration \(y\), in hours, for a set of teens. For a screen-time value of 2 hours, the fitted line predicts 7.4 hours of sleep, and the residual standard deviation is \(s=0.8\) hour. Compare the prediction for one teen with the average for 16 independent teens with that same screen-time value.

State. The line’s predicted response at 2 hours is 7.4 hours. That value is a predicted mean response for teens at that screen-time level, and it can also be used as a point prediction for one teen. The individual response is less precise because individual teens vary around the predicted mean.

Plan. Describe the individual’s typical residual scale using \(s=0.8\) hour. Then apply the averaging comparison to 16 independent teens, assuming their responses have similar scatter and the same model prediction.

Do. For one teen, the typical residual scale is 0.8 hour. For the average of 16 teens:

$$ \frac{s}{\sqrt{m}} =\frac{0.8}{\sqrt{16}} =\frac{0.8}{4} =0.2\text{ hour} $$

The predicted average remains 7.4 hours of sleep. The approximate scatter of the group average around the line is 0.2 hour, one quarter of the individual residual scale.

Conclude in context. The regression line predicts an average of 7.4 hours of sleep for teens with 2 hours of evening screen time. For one teen, the typical residual scale is about 0.8 hour. For an average of 16 independent, comparable teens at that screen-time value, the approximate scale is 0.2 hour. Neither number specifies the exact sleep duration of a particular teen or guarantees how close the group average will be to 7.4 hours.

Worked Example: A Group-Level Model Does Not Predict Each Student

Hypothetical setting. A school district fits a regression using one point per school: \(x\) is the average weekly tutoring time per student, and \(y\) is the school’s average quiz score. At one tutoring level, the line predicts a school mean score of 78 points, and the residual standard deviation for this group-level model is 4 points. The district asks whether it can use 4 points as the typical prediction error for an individual student.

State. The regression’s cases are schools, not students. Therefore, its fitted value of 78 points predicts a school average quiz score, and its \(s=4\) points describes the typical scatter of observed school averages around the fitted line. It does not describe how much individual students’ scores vary within a school.

Plan. Identify what one row of the regression data represents, then match the meaning of \(s\) to those cases. To compare an individual student with a class average, use a separate measure of within-school student variation if one is supplied; do not substitute the group-level model’s \(s\).

Do. Suppose the standard deviation of individual student scores within schools at this tutoring level is about 12 points. The group-level model’s \(s=4\) points is the residual scale for school averages. If 9 independent students are sampled from a comparable school, the approximate scatter of their sample average around that school’s mean, based on the stated within-school variation, is:

$$ \frac{12}{\sqrt{9}} =\frac{12}{3} =4\text{ points} $$

That 4-point calculation comes from the within-school student variation and the sample size—not from the regression’s group-level residual standard deviation, even though the two values happen to match in this example.

Conclude in context. The group-level regression predicts a school average of 78 points; its \(s=4\) points describes scatter among school averages around the line. It cannot be used as the typical error for predicting one student’s quiz score. The 12-point within-school standard deviation describes individual variation, and the approximate 4-point scatter for an average of 9 students is based on that different source of variation.

Common Mistakes and AP Exam Tips

  • Assuming the fitted value is what every individual will get. A regression line predicts the center or mean response at \(x\). State that individual responses vary around the line.
  • Calling \(s\) a guaranteed distance. The residual standard deviation describes a typical scale of scatter. It is not a maximum error, and it does not say that every response falls within \(s\) of its fitted value.
  • Applying \(s/\sqrt{m}\) without checking independence. Averaging is less effective when responses are related. Mention independence and comparable scatter when using this approximate rule.
  • Treating \(s/\sqrt{m}\) as an interval or probability. This calculation gives an approximate scale for average residual variation. By itself, it does not provide a confidence level or the probability that an average falls within a chosen distance.
  • Using group-level residual scatter for individuals. As emphasized in “Group Averages Versus Individual Predictions,” identify what counts as a case in the regression. A model fitted to group means describes group means, not the individual observations within each group.
  • Confusing the predicted mean with the prediction’s precision. The predicted mean can stay the same for one individual and for a group average, while the average is more precise because individual deviations can offset one another.

A strong AP-style response names what the model predicts and distinguishes the cases: “At this watering level, the line predicts an average plant height of 42 centimeters. For one plant, responses typically scatter around the line on a scale of about 6 centimeters. For the average of 9 independent, comparable plants, the approximate scatter scale is 2 centimeters. This is not a guarantee for an individual plant or for the group average.”

Key takeaway: A fitted value can represent a predicted mean, but one individual’s response may vary substantially around it. Averaging independent, comparable responses tends to reduce scatter; a group-level prediction or residual standard deviation should not be mistaken for an individual prediction.

Check Your Understanding

Use the distinction between individual predictions and mean predictions to answer each question.

  1. A model predicts a mean response of 30 units at a given \(x\), with \(s=5\) units. What does the prediction say about one individual, and what does \(s\) say?
  2. If \(s=8\) minutes, what is the approximate residual-scatter scale for the average of 16 independent, comparable responses at the same \(x\)?
  3. Why might averaging repeated measurements from the same person fail to reduce scatter as much as averaging independent people?
  4. A regression uses one point per neighborhood, where each response is the neighborhood’s average travel time. Can its residual standard deviation be used as the typical error for predicting one resident’s travel time? Explain.
  5. What does the calculation \(s/\sqrt{m}\) describe, and what does it not guarantee?