Adding a Group to a Scatterplot
In “Comparing Two Scatterplots,” you learned to compare form and strength when separate plots show the same variables. A single scatterplot can make that comparison more direct by showing a third variable: a categorical variable that identifies a group for each observational unit. The horizontal and vertical axes still display the same two quantitative variables; color or plotting symbols distinguish the categories.
For example, a scatterplot might show the relationship between weekly practice time and a performance score, with students’ music programs identified by different symbols. Each point still represents one student and that student’s paired values. The symbol adds information about which program the student belongs to; it does not represent a third numerical measurement.
The added category can reveal whether the pattern looks similar across groups or differs in direction, form, strength, or location. It may also reveal that an overall pattern is partly due to the groups having different ranges or typical response values. To interpret the plot carefully, first read the axes and legend, then describe the pattern for each group, and finally compare the groups.
This is a visual extension of DUFS, introduced in “Describing a Scatterplot With DUFS.” Direction, unusual features, form, and strength still apply, but you should be clear about whether you mean the pattern within a particular group or the pattern formed by all points together. Those are not always the same description.
How to Read the Group Coding
Begin with the individuals and variables. Each point represents one observational unit, such as a student, plant, or device. The horizontal axis displays the explanatory variable, and the vertical axis displays the response variable, as in “Explanatory and Response Variables in Scatterplots.” The categorical variable identifies a group for each point.
Next, use the legend to identify the categories. If a plot uses blue circles for one group and orange triangles for another, do not rely on color alone: confirm the legend’s labels and match them to the plotting symbols. A well-designed plot makes categories distinct, including for readers who may not distinguish colors easily.
Name the two quantitative variables, their units, and which is explanatory and which is response.
Identify the category represented by each color or symbol. Check whether every point belongs to one of the displayed groups.
For each category, consider direction, unusual features, form, and strength. Do not assume all categories share one pattern.
State how their patterns are alike or different, including any differences in location or range that matter to the display.
Notice what all points together suggest, but do not use that pattern as a substitute for examining the groups separately.
Keep the earlier guidance about comparable scales in mind. When comparing group patterns in one plot, all categories share the same axes, which makes their displayed locations and spreads directly comparable. Still, a group with a different range of explanatory-variable values may show a different portion of a curved relationship, so interpret the visible patterns in light of the data range.
Compare Direction, Form, and Strength Within Groups
Describe each group’s association on its own terms. A useful description might say that both groups have positive, roughly linear patterns, but one group’s points lie more closely around its trend. Or the groups might differ in form: one pattern could be roughly linear while the other bends. The same principles from “Judging Strength of an Association” and “Association Versus Linear Association” apply: form and strength are distinct.
Also note differences in location. One group’s points may generally lie above another group’s points across a shared range of explanatory values. That is a visual description of the plotted response values, not evidence by itself about why the groups differ. Similarly, a steeper-looking pattern is not automatically stronger; strength concerns how closely points follow the pattern within that group.
Worked Example: Comparing Two Training Groups
Imagine a fictional program recording practice hours and a skill score for students in two training groups. Each student is one point. Circles identify Group A and triangles identify Group B. The table gives invented observations, not results from a real study.
| Practice hours | Group A score (points) | Group B score (points) |
|---|---|---|
| 1 | 22 | 15 |
| 2 | 25 | 21 |
| 3 | 29 | 26 |
| 4 | 31 | 32 |
| 5 | 36 | 36 |
| 6 | 38 | 43 |
Read the display. Practice hours is the explanatory variable on the horizontal axis, skill score in points is the response variable on the vertical axis, and the legend distinguishes the two groups.
Describe each pattern. In both groups, higher practice-hour values tend to go with higher scores. Both groups show a positive, roughly linear pattern. The points in each group stay fairly close to an upward trend; the displayed observations do not suggest a major difference in form.
Compare location and strength. At the listed practice-hour values, Group A’s scores are generally somewhat higher than Group B’s at the lower hours, while the scores are similar near five hours. Group B’s points also appear to follow a positive trend, though the visible scatter is not identical. With only these points, describe any strength difference cautiously rather than claiming a precise or conclusive ranking.
Conclusion in context. In these invented observations, practice hours and skill score have positive, roughly linear associations in both training groups. The groups’ patterns are broadly similar, although their scores differ somewhat at some practice-hour values. The plot describes an association; it does not establish that practice caused the score differences or that group membership explains them.
Why the Combined Pattern Can Be Misleading
A common first glance is to ignore the symbols and describe all points together. That can hide important differences. For example, two groups could each have a weak positive pattern, but one group may tend to have higher response values and also occupy larger explanatory values. The pooled points could then show a strong upward pattern, even though the within-group patterns are much weaker. In another situation, groups could show different within-group directions, while the pooled pattern looks quite different from either one.
This does not mean that the combined pattern is always useless. It answers a different descriptive question: what pattern do all observations show when group labels are ignored? The group-specific patterns answer how the two quantitative variables are associated within each category. When the plot includes a group code, inspect both levels and state clearly which one you are describing.
Worked Example: A Pooled Trend That Hides Group Differences
A fictional nature club records the number of hours volunteers spend outdoors each month and a nature-identification score. Members belong to either the Morning group or the Evening group. Symbols distinguish the groups. The invented points below are chosen to illustrate why the labels matter.
| Group | Outdoor hours | Identification score (points) |
|---|---|---|
| Morning | 1 | 31 |
| Morning | 2 | 34 |
| Morning | 3 | 32 |
| Evening | 5 | 18 |
| Evening | 6 | 21 |
| Evening | 7 | 19 |
Look within the Morning group. The scores stay around 31 to 34 points as outdoor hours increase from 1 to 3. There is little visible association in these three observations; they do not show a clear upward pattern.
Look within the Evening group. Scores stay around 18 to 21 points as outdoor hours increase from 5 to 7. This group also shows little visible association in these observations.
Now consider all points together. The Morning observations have both lower outdoor-hour values and higher scores, while the Evening observations have higher outdoor-hour values and lower scores. If the group labels are ignored, the combined cloud may appear to have a negative association: larger outdoor-hour values tend to occur with lower scores. That overall impression mainly reflects the separation between the groups, not a clear negative trend within either group.
Conclusion in context. In these invented observations, neither group shows a strong within-group association between outdoor hours and identification score, but the groups differ in where their points are located. Combining them gives a negative-looking pattern that would be misleading if described as a general within-group relationship. The plot alone does not explain the group difference.
The lesson is not to discard the combined pattern, but to report it at the correct level. A precise description might say, “Together, the points show a negative-looking pattern, while within each group the scores vary little across the observed hours; the groups occupy different regions of the plot.” This makes both observations clear without treating the pooled trend as the story for every group.
When Groups Differ in Form
Group coding can also reveal that groups have different forms. One category might show a roughly linear increase, while another shows a curve. If the groups cover different ranges of the explanatory variable, be careful: a curve can look almost straight over a short range. Describe the visible pattern and its range rather than assuming that an apparent difference must hold beyond the data shown.
Worked Example: Linear and Curved Patterns by Equipment Type
A fictional lab tracks operating time and cooling output for devices using two equipment types. The horizontal axis is operating time in hours; the vertical axis is cooling output in units per minute. Squares and diamonds identify equipment type. These invented values illustrate differences in form.
| Operating time (hours) | Type A output (units/minute) | Type B output (units/minute) |
|---|---|---|
| 0 | 5 | 4 |
| 1 | 8 | 11 |
| 2 | 11 | 16 |
| 3 | 14 | 19 |
| 4 | 17 | 21 |
Describe Type A. Its output rises by about 3 units per minute for each additional hour in the table, making the pattern positive and roughly linear.
Describe Type B. Output increases as operating time increases, but the increases become smaller: 7, 5, 3, then 2 units per minute. The pattern rises and begins to level off, so it is positive and curved rather than roughly linear.
Compare the groups. Both types have positive associations between operating time and cooling output, but their forms differ: Type A is roughly linear, while Type B bends and begins to level off. The table offers less basis for a strong comparison of scatter because the values follow their respective patterns closely and there are only five observations per type.
Conclusion in context. For these invented device observations, output tends to increase with operating time for both equipment types. The increase appears roughly steady for Type A and smaller at later times for Type B. These patterns describe the plotted devices and do not establish a causal explanation or predict behavior outside the displayed time range.
Common Mistakes and AP Exam Tips
- Ignoring the legend: A color or symbol has no useful group meaning until you connect it to the legend. Name the group in your description, not just “the blue points.”
- Describing only the pooled cloud: The overall pattern may differ from the patterns within groups. Inspect and describe each category before summarizing all points together.
- Confusing group location with strength: One group may lie higher on the graph but have more scatter around its own pattern. Vertical position and strength are different features.
- Calling the steepest group the strongest: Steepness concerns how quickly the response changes across the explanatory-variable scale. Strength concerns how closely points follow their group’s pattern.
- Treating the categorical variable as a numerical axis: In a group-coded scatterplot, the category is represented by symbols or colors; the two axes still show the quantitative variables.
- Claiming cause and effect: The display shows how variables and group labels occur together. It does not, by itself, show that category membership caused a difference in the response.
- Overstating a pattern from a few observations: Describe what the plotted points suggest, and be cautious about fine distinctions in form or strength when there are few points or substantial overlap.
For full credit, name the quantitative variables and units, identify the groups, and state how the patterns compare. A strong response does more than say “the groups are different.” It specifies whether the difference concerns direction, form, strength, or location, and it avoids claiming more than the plot supports.
Check Your Understanding
Use the axes, legend, and group-specific patterns to answer each question.
- A scatterplot shows study hours on the horizontal axis, exam score on the vertical axis, and circles and triangles for two school programs. What does each point represent, and what does the symbol add?
- Both groups have positive, roughly linear patterns, but one group’s points lie more closely around its own trend. Which group has the stronger association, and what feature supports that comparison?
- Why should you describe group-specific patterns before describing the combined cloud?
- One group’s points lie higher than another group’s at similar explanatory-variable values, but both groups have similar scatter. Which feature differs, and which feature may be similar?
- Write a contextual sentence comparing a roughly linear pattern in one category with a positive, curved pattern that levels off in another. Avoid making a causal claim.