Tutorials › AP Statistics › Group Averages Versus Individual Predictions

Regression and context · Tutorial 933 of 1000

Group Averages Versus Individual Predictions

Distinguish predictions about group averages from predictions about individual cases, even when a regression fitted to group means appears to fit extremely well.

Intermediate 12 min read

What You'll Learn

  • Identify whether the cases in a regression are individuals or groups.
  • Explain why averaging can hide individual variation.
  • Interpret a regression line, residual, \(r^2\), and \(s\) at the level of the data used to fit the model.
  • Avoid treating a relationship between group averages as an individual-level relationship.
  • State what additional data are needed to evaluate predictions for individuals.

Averages Describe Groups, Not Automatically Their Members

A regression equation can fit a set of group averages very closely and still be a poor guide to any one person in those groups. This is a different issue from the sampling concerns in “Sampling Method and Reliability of a Model”: even when the groups are relevant to the question, the level at which the data were recorded matters. A point in the regression might represent one person—or it might represent a school, clinic, neighborhood, or other group.

Suppose a researcher records each school’s average weekly tutoring time and average exam score. A regression using those values describes how schools’ averages vary together. Its predictions concern the average exam score for a school with a given average tutoring time. It does not, by itself, predict an individual student’s exam score from that student’s tutoring time.

Definition: A group-level regression uses groups as its cases, often with one point per group. A prediction from that model describes the response average for a group with the stated explanatory-variable average. An individual-level prediction concerns one person or object and requires evidence about individual cases.

Averages smooth over differences among the people or objects inside each group. A school’s average score does not show whether nearly all students scored close to that average or whether some scored much lower and others much higher. When the model uses only school averages, that within-school variation is not represented by the points in its scatterplot.

The distinction affects every regression summary. As in “Comparing Residuals Across Observations,” a residual is observed minus predicted, but the relevant “observation” is the case used to fit the model. For a group-level model, the residual compares a group’s observed average response with its predicted average response. Likewise, \(r^2\) describes variation in the group averages, and \(s\) describes residual scatter for those group-level cases. Neither automatically measures how accurately the model predicts individuals.

Key distinction: A small \(s\) or high \(r^2\) for group averages is evidence about how the model fits those averages. It is not a measure of individual prediction accuracy. Always identify what one data point represents before interpreting the model.

Why a Strong Fit to Averages Can Mislead

When a group average is calculated, individual highs and lows can balance each other. The group’s average may follow a smooth pattern across groups even though the individuals within any one group differ considerably. A regression of averages can therefore have small residuals while individuals’ responses are far from their group’s predicted average.

This does not mean that averaging always produces a high \(r^2\), a small \(s\), or a misleading model. The result depends on the data. The important point is that group-level fit statistics answer a group-level question. They do not supply information that was not measured, such as how individual responses vary with individual explanatory-variable values.

There is another risk: a relationship between group averages need not match the relationship among individuals. For example, schools with more average study time might also have higher average scores, but that alone does not show that a student who studies more than another student at the same school will score higher. Students within a school may differ in many other ways, and the group averages do not reveal the paired individual data needed to examine that question.

Interpretation check: Before saying “the model predicts a person’s response,” ask: Were the model’s cases people, or were they groups of people? If the cases were groups, describe the prediction as a predicted group average unless individual-level evidence is also available.

This careful wording avoids an ecological inference: drawing a conclusion about individuals from a pattern observed across groups. The name is less important than the reasoning. A school-level pattern is evidence about schools; it cannot simply be assigned to every student in those schools.

Worked Examples: Keeping the Level of Prediction Clear

Worked Example: Tutoring and Exam Scores Across Schools

A fictional district records the average hours of tutoring per week, \(x\), and average exam score, \(y\), for four schools. The values are shown below. Each row is one school, not one student.

SchoolAverage tutoring hours, \(x\)Average exam score, \(y\)
A250
B460
C670
D880

State. Describe what the fitted line predicts, and decide whether a perfect fit to these four points establishes accurate predictions for individual students.

Plan. First identify the cases. Because each point is a school average, calculate and interpret the regression at the school level. Then compare a predicted average with some individual scores from one school.

Do. The four points lie exactly on the line \(\hat{y}=40+5x\). For example, at \(x=4\), the fitted value is

$$ \hat{y}=40+5(4)=60. $$

That is a predicted average exam score of 60 for a school whose average tutoring time is 4 hours per week. For these four group-level points, every residual is 0. Thus \(SSE=0\), \(r^2=1\), and, using \(s=\sqrt{SSE/(n-2)}\) with \(n=4\) schools, \(s=\sqrt{0/(4-2)}=0\) score points. These summaries describe an exact fit to the four school averages.

Now suppose the four students in School B have scores 70, 50, 60, and 60. Their average is \((70+50+60+60)/4=60\), matching the school average. The fitted value of 60 is right for the school’s average, but it is not a close prediction for the student who scored 70 or the student who scored 50. The group-level regression has no individual student points with which to assess those differences.

Conclude. “The line fits the four school averages exactly and predicts an average score for a school at a given average tutoring time. It does not establish that individual students’ scores can be predicted accurately from their own tutoring time.”

Worked Example: Reading Time and Class Averages

A fictional teacher compares four classes. For each class, \(x\) is the class’s average minutes of independent reading per day and \(y\) is the class’s average quiz score, in points. The group-average pairs are \((2,52)\), \((4,58)\), \((6,73)\), and \((8,77)\). Find and interpret a regression model for these class averages, then consider whether it predicts an individual student’s score.

Do the group-level calculation. The means of the four \(x\)-values and four \(y\)-values are \(\bar{x}=5\) minutes and \(\bar{y}=65\) points. The sum of the cross-products of deviations is

$$ (-3)(-13)+(-1)(-7)+(1)(8)+(3)(12)=90. $$

The sum of squared \(x\)-deviations is \(9+1+1+9=20\), so the slope is \(90/20=4.5\) points per minute of class-average reading time. The intercept is \(65-4.5(5)=42.5\) points. The fitted line is \(\hat{y}=42.5+4.5x\).

For the four classes, the predicted averages are 51.5, 60.5, 69.5, and 78.5 points. The residuals, observed minus predicted, are \(0.5\), \(-2.5\), \(3.5\), and \(-1.5\) points. Therefore, \(SSE=0.5^2+(-2.5)^2+3.5^2+(-1.5)^2=21\). The total variation in the four class averages is \(SST=(-13)^2+(-7)^2+8^2+12^2=426\). Thus

$$ r^2=1-\frac{SSE}{SST}=1-\frac{21}{426}\approx 0.9507. $$

The residual standard deviation for these four class-level cases is

$$ s=\sqrt{\frac{SSE}{n-2}}=\sqrt{\frac{21}{4-2}}=\sqrt{10.5}\approx 3.24\text{ points}. $$

The \(r^2\) value describes the fraction of variation in these class average quiz scores accounted for by a linear relationship with class average reading time. The \(s\) value describes the typical residual size around the line for class averages, in quiz-score points. Neither gives the typical prediction error for individual students.

For the class with average reading time 4 minutes, the predicted class average is 60.5 points. Suppose that class’s four quiz scores are 43, 53, 63, and 73; their average is 58 points, so the group residual is \(58-60.5=-2.5\) points. If 60.5 were mistakenly used as a prediction for each of those students, the differences between observed scores and that value would be \(-17.5\), \(-7.5\), \(2.5\), and \(12.5\) points. Those differences illustrate individual variation; they are not the class-level residuals used to calculate \(s\) above.

Conclude. “The line describes class averages well: about 95.07% of the variation in the four class average quiz scores is accounted for by their linear relationship with class average reading time. It does not show that the model predicts individual students’ quiz scores with a typical error of 3.24 points.”

Worked Example: Average Wait and Patient Ratings Across Clinics

A fictional health network compares three clinics. Each point represents one clinic’s average patient wait, \(x\), in minutes, and its mean satisfaction rating, \(y\), on a 0-to-10 scale. The group averages are \((10,8)\), \((20,7)\), and \((30,6)\). A manager proposes using the fitted line to predict a particular patient’s satisfaction rating when that patient waits 20 minutes.

State. Decide whether that individual prediction is supported by the regression on clinic averages.

Plan. Use the four-step regression reasoning: identify the cases and target, check what the fitted model predicts, distinguish a clinic average from an individual response, and make a conclusion limited to what the data support.

Do. The three clinic-average points lie on \(\hat{y}=9-0.1x\). For a clinic with an average wait of 20 minutes, the fitted satisfaction rating is

$$ \hat{y}=9-0.1(20)=7. $$

This is a predicted clinic mean of 7 rating points for a clinic whose average wait is 20 minutes. The model’s cases are clinics, not patients. A particular patient’s 20-minute wait is not the same thing as a clinic-wide average wait of 20 minutes, and the clinic’s mean rating does not reveal how ratings vary among its patients.

For instance, four patients at a clinic might give ratings of 4, 6, 8, and 10, whose mean is \((4+6+8+10)/4=7\). Knowing that the clinic mean is 7 does not tell us which rating any one patient gave. Nor do the three clinic averages provide paired patient-level wait times and satisfaction ratings from which to assess an individual relationship.

Conclude. “The group-level line predicts a mean satisfaction rating of 7 for a clinic with a 20-minute average wait. It does not support predicting a particular patient’s rating from that patient’s wait; individual-level paired data would be needed to evaluate that prediction.”

Common Mistakes and AP Exam Tips

The most important step is to name the observational unit: what does each point represent? A response that correctly interprets the calculation but applies it to the wrong level can still make an unsupported claim.

  • Calling a group prediction an individual prediction. Say “predicted average score for a school” when the model’s cases are schools. Do not replace “school” with “student” unless the model was fitted to individual student data.
  • Applying \(r^2\) to the wrong responses. If the points are clinics, \(r^2\) describes variation in clinic averages, not the percentage of variation in individual patients’ ratings.
  • Giving \(s\) the wrong meaning. As in “Standard Deviation of the Residuals, \(s\),” state its response units and typical residual size. For a model fitted to group means, those are group-level residuals—not individual prediction errors.
  • Assuming a perfect or nearly perfect group fit proves individual accuracy. Explain that within-group variation is not represented by a scatterplot of group averages. A strong group pattern cannot substitute for individual-level data.
  • Claiming group and individual relationships must be the same. The group-average pattern does not establish the pattern among individuals. Avoid conclusions about what happens to one person based only on comparisons across groups.
  • Forgetting what one point represents when counting cases. In a model with one point per school, \(n\) is the number of schools in that regression, not the total number of students attending them.

A full-credit AP response makes the unit of analysis explicit, interprets the model summary for those cases, and states why that result does or does not answer an individual-level question. If the question asks for an individual prediction but supplies only group averages, say that the individual prediction is not established by the given regression.

Key takeaway: A regression fitted to means predicts means. It may fit those group averages closely while individuals within each group vary substantially. Interpret the line, residuals, \(r^2\), and \(s\) at the level of the cases used to fit the model; individual predictions require individual-level evidence.

Check Your Understanding

For each situation, identify what one point represents and keep the model’s prediction at that same level.

  1. A regression uses one point per town, with average daily bicycle trips and average air-quality index. What does a fitted value predict, and what would it not establish about one resident?
  2. A model of school averages has \(r^2=0.88\). State what variation this value describes and name one thing it does not tell you about students.
  3. A clinic-level model has \(s=1.2\) satisfaction-rating points. Explain what residuals this value summarizes and why it is not automatically an individual patient’s typical prediction error.
  4. In your own words, explain how individual scores can vary considerably even when group means fall close to a straight line.
  5. What kind of data would be needed to assess whether a person’s own explanatory-variable value predicts that person’s response?