Averages Describe Groups, Not Automatically Their Members
A regression equation can fit a set of group averages very closely and still be a poor guide to any one person in those groups. This is a different issue from the sampling concerns in “Sampling Method and Reliability of a Model”: even when the groups are relevant to the question, the level at which the data were recorded matters. A point in the regression might represent one person—or it might represent a school, clinic, neighborhood, or other group.
Suppose a researcher records each school’s average weekly tutoring time and average exam score. A regression using those values describes how schools’ averages vary together. Its predictions concern the average exam score for a school with a given average tutoring time. It does not, by itself, predict an individual student’s exam score from that student’s tutoring time.
Averages smooth over differences among the people or objects inside each group. A school’s average score does not show whether nearly all students scored close to that average or whether some scored much lower and others much higher. When the model uses only school averages, that within-school variation is not represented by the points in its scatterplot.
The distinction affects every regression summary. As in “Comparing Residuals Across Observations,” a residual is observed minus predicted, but the relevant “observation” is the case used to fit the model. For a group-level model, the residual compares a group’s observed average response with its predicted average response. Likewise, \(r^2\) describes variation in the group averages, and \(s\) describes residual scatter for those group-level cases. Neither automatically measures how accurately the model predicts individuals.
Why a Strong Fit to Averages Can Mislead
When a group average is calculated, individual highs and lows can balance each other. The group’s average may follow a smooth pattern across groups even though the individuals within any one group differ considerably. A regression of averages can therefore have small residuals while individuals’ responses are far from their group’s predicted average.
This does not mean that averaging always produces a high \(r^2\), a small \(s\), or a misleading model. The result depends on the data. The important point is that group-level fit statistics answer a group-level question. They do not supply information that was not measured, such as how individual responses vary with individual explanatory-variable values.
There is another risk: a relationship between group averages need not match the relationship among individuals. For example, schools with more average study time might also have higher average scores, but that alone does not show that a student who studies more than another student at the same school will score higher. Students within a school may differ in many other ways, and the group averages do not reveal the paired individual data needed to examine that question.
This careful wording avoids an ecological inference: drawing a conclusion about individuals from a pattern observed across groups. The name is less important than the reasoning. A school-level pattern is evidence about schools; it cannot simply be assigned to every student in those schools.
Worked Examples: Keeping the Level of Prediction Clear
Worked Example: Tutoring and Exam Scores Across Schools
A fictional district records the average hours of tutoring per week, \(x\), and average exam score, \(y\), for four schools. The values are shown below. Each row is one school, not one student.
| School | Average tutoring hours, \(x\) | Average exam score, \(y\) |
|---|---|---|
| A | 2 | 50 |
| B | 4 | 60 |
| C | 6 | 70 |
| D | 8 | 80 |
State. Describe what the fitted line predicts, and decide whether a perfect fit to these four points establishes accurate predictions for individual students.
Plan. First identify the cases. Because each point is a school average, calculate and interpret the regression at the school level. Then compare a predicted average with some individual scores from one school.
Do. The four points lie exactly on the line \(\hat{y}=40+5x\). For example, at \(x=4\), the fitted value is
That is a predicted average exam score of 60 for a school whose average tutoring time is 4 hours per week. For these four group-level points, every residual is 0. Thus \(SSE=0\), \(r^2=1\), and, using \(s=\sqrt{SSE/(n-2)}\) with \(n=4\) schools, \(s=\sqrt{0/(4-2)}=0\) score points. These summaries describe an exact fit to the four school averages.
Now suppose the four students in School B have scores 70, 50, 60, and 60. Their average is \((70+50+60+60)/4=60\), matching the school average. The fitted value of 60 is right for the school’s average, but it is not a close prediction for the student who scored 70 or the student who scored 50. The group-level regression has no individual student points with which to assess those differences.
Conclude. “The line fits the four school averages exactly and predicts an average score for a school at a given average tutoring time. It does not establish that individual students’ scores can be predicted accurately from their own tutoring time.”
Worked Example: Reading Time and Class Averages
A fictional teacher compares four classes. For each class, \(x\) is the class’s average minutes of independent reading per day and \(y\) is the class’s average quiz score, in points. The group-average pairs are \((2,52)\), \((4,58)\), \((6,73)\), and \((8,77)\). Find and interpret a regression model for these class averages, then consider whether it predicts an individual student’s score.
Do the group-level calculation. The means of the four \(x\)-values and four \(y\)-values are \(\bar{x}=5\) minutes and \(\bar{y}=65\) points. The sum of the cross-products of deviations is
The sum of squared \(x\)-deviations is \(9+1+1+9=20\), so the slope is \(90/20=4.5\) points per minute of class-average reading time. The intercept is \(65-4.5(5)=42.5\) points. The fitted line is \(\hat{y}=42.5+4.5x\).
For the four classes, the predicted averages are 51.5, 60.5, 69.5, and 78.5 points. The residuals, observed minus predicted, are \(0.5\), \(-2.5\), \(3.5\), and \(-1.5\) points. Therefore, \(SSE=0.5^2+(-2.5)^2+3.5^2+(-1.5)^2=21\). The total variation in the four class averages is \(SST=(-13)^2+(-7)^2+8^2+12^2=426\). Thus
The residual standard deviation for these four class-level cases is
The \(r^2\) value describes the fraction of variation in these class average quiz scores accounted for by a linear relationship with class average reading time. The \(s\) value describes the typical residual size around the line for class averages, in quiz-score points. Neither gives the typical prediction error for individual students.
For the class with average reading time 4 minutes, the predicted class average is 60.5 points. Suppose that class’s four quiz scores are 43, 53, 63, and 73; their average is 58 points, so the group residual is \(58-60.5=-2.5\) points. If 60.5 were mistakenly used as a prediction for each of those students, the differences between observed scores and that value would be \(-17.5\), \(-7.5\), \(2.5\), and \(12.5\) points. Those differences illustrate individual variation; they are not the class-level residuals used to calculate \(s\) above.
Conclude. “The line describes class averages well: about 95.07% of the variation in the four class average quiz scores is accounted for by their linear relationship with class average reading time. It does not show that the model predicts individual students’ quiz scores with a typical error of 3.24 points.”
Worked Example: Average Wait and Patient Ratings Across Clinics
A fictional health network compares three clinics. Each point represents one clinic’s average patient wait, \(x\), in minutes, and its mean satisfaction rating, \(y\), on a 0-to-10 scale. The group averages are \((10,8)\), \((20,7)\), and \((30,6)\). A manager proposes using the fitted line to predict a particular patient’s satisfaction rating when that patient waits 20 minutes.
State. Decide whether that individual prediction is supported by the regression on clinic averages.
Plan. Use the four-step regression reasoning: identify the cases and target, check what the fitted model predicts, distinguish a clinic average from an individual response, and make a conclusion limited to what the data support.
Do. The three clinic-average points lie on \(\hat{y}=9-0.1x\). For a clinic with an average wait of 20 minutes, the fitted satisfaction rating is
This is a predicted clinic mean of 7 rating points for a clinic whose average wait is 20 minutes. The model’s cases are clinics, not patients. A particular patient’s 20-minute wait is not the same thing as a clinic-wide average wait of 20 minutes, and the clinic’s mean rating does not reveal how ratings vary among its patients.
For instance, four patients at a clinic might give ratings of 4, 6, 8, and 10, whose mean is \((4+6+8+10)/4=7\). Knowing that the clinic mean is 7 does not tell us which rating any one patient gave. Nor do the three clinic averages provide paired patient-level wait times and satisfaction ratings from which to assess an individual relationship.
Conclude. “The group-level line predicts a mean satisfaction rating of 7 for a clinic with a 20-minute average wait. It does not support predicting a particular patient’s rating from that patient’s wait; individual-level paired data would be needed to evaluate that prediction.”
Common Mistakes and AP Exam Tips
The most important step is to name the observational unit: what does each point represent? A response that correctly interprets the calculation but applies it to the wrong level can still make an unsupported claim.
- Calling a group prediction an individual prediction. Say “predicted average score for a school” when the model’s cases are schools. Do not replace “school” with “student” unless the model was fitted to individual student data.
- Applying \(r^2\) to the wrong responses. If the points are clinics, \(r^2\) describes variation in clinic averages, not the percentage of variation in individual patients’ ratings.
- Giving \(s\) the wrong meaning. As in “Standard Deviation of the Residuals, \(s\),” state its response units and typical residual size. For a model fitted to group means, those are group-level residuals—not individual prediction errors.
- Assuming a perfect or nearly perfect group fit proves individual accuracy. Explain that within-group variation is not represented by a scatterplot of group averages. A strong group pattern cannot substitute for individual-level data.
- Claiming group and individual relationships must be the same. The group-average pattern does not establish the pattern among individuals. Avoid conclusions about what happens to one person based only on comparisons across groups.
- Forgetting what one point represents when counting cases. In a model with one point per school, \(n\) is the number of schools in that regression, not the total number of students attending them.
A full-credit AP response makes the unit of analysis explicit, interprets the model summary for those cases, and states why that result does or does not answer an individual-level question. If the question asks for an individual prediction but supplies only group averages, say that the individual prediction is not established by the given regression.
Check Your Understanding
For each situation, identify what one point represents and keep the model’s prediction at that same level.
- A regression uses one point per town, with average daily bicycle trips and average air-quality index. What does a fitted value predict, and what would it not establish about one resident?
- A model of school averages has \(r^2=0.88\). State what variation this value describes and name one thing it does not tell you about students.
- A clinic-level model has \(s=1.2\) satisfaction-rating points. Explain what residuals this value summarizes and why it is not automatically an individual patient’s typical prediction error.
- In your own words, explain how individual scores can vary considerably even when group means fall close to a straight line.
- What kind of data would be needed to assess whether a person’s own explanatory-variable value predicts that person’s response?