Tutorials › AP Statistics › Identifying Clusters and Gaps

Scatterplots and association · Tutorial 808 of 1000

Identifying Clusters and Gaps

Learn how clusters and gaps can reveal subgroups or limits in the data, and how to describe these patterns carefully in context.

Intermediate 9 min read

What You'll Learn

  • Distinguish a cluster, a gap, and an isolated point in a scatterplot
  • Describe where clusters and gaps occur using both explanatory and response variables
  • Use labels and study context to consider possible subgroup explanations
  • Explain why a visible gap does not, by itself, prove a cause or a population boundary
  • Avoid treating every sparse region as an outlier or assuming one overall pattern describes every subgroup

Look for Groups as Well as Individual Points

In “Spotting Outliers in Bivariate Data,” you learned to judge an unusual point in relation to the overall pattern. Now widen your view: a scatterplot can contain several groups of points, or an empty region separating parts of the plot. These features may be clues about how the data were collected or about meaningful subgroups in the observations.

A cluster is a concentration of points in one region of a scatterplot. A gap is a region with few or no points separating parts of the display. A gap may be visible horizontally, vertically, or diagonally between clouds of points. To describe either feature, refer to both variables and their units where possible; saying only that “the points are split” leaves the reader unsure where.

Definition: A cluster is a group of observations concentrated in one area of a scatterplot. A gap is an area with few or no observations between regions of the plot. Clusters and gaps describe the arrangement of paired \((x,y)\) values; context and available labels help you investigate what they might indicate.

A scatterplot does not explain why a cluster or gap exists. Possible explanations include distinct subgroups, different settings or conditions, how the sample was selected, or a range of values that was not observed. Those are possibilities to investigate, not conclusions established by the picture alone. Check the study context and any group labels before proposing an explanation.

Also distinguish a cluster from an outlier. A cluster contains multiple observations concentrated together; an outlier is an individual point far from the overall pattern, as discussed in the previous tutorial. A point by itself in a gap may deserve attention, but its position alone does not tell you whether it is an error, a valid unusual observation, or part of an unrecognized subgroup.

A Careful Way to Describe Clusters and Gaps

First identify where the denser groups are and where the sparse or empty regions lie. Then describe their locations using the variables: for example, a group might have smaller \(x\)-values and lower \(y\)-values than another group. If the graph includes labels, compare the visible groups with those labels. Finally, state what the pattern might indicate while making clear that the plot alone cannot establish a cause.

1
Locate the dense regions.
Look for areas where several points lie relatively close together. Note each region’s approximate \(x\)- and \(y\)-values rather than relying only on labels such as “left” or “high.”
2
Find any sparse or empty areas.
Check whether a region with few or no points separates the clusters. State which values of \(x\), \(y\), or both characterize that region.
3
Compare with labels and context.
Ask whether observations in different clusters share a known feature, such as location, group, time period, or setting. Do not assign a label the plot does not provide.
4
Interpret cautiously.
Explain what the arrangement may suggest, then note that a scatterplot alone cannot confirm why the groups or gap occur.

The “overall pattern” can be misleading if distinct clusters are treated as one undifferentiated cloud. A relationship visible within one group may not describe another group, and the full set of points may look different from either group considered separately. For this tutorial, the key skill is to notice that structure and describe it—not to assume that one single pattern tells the whole story.

A gap is relative to the data shown. It means that no, or very few, observed points occupy a region of this particular scatterplot. It does not mean that the values in that region are impossible, that the population has a sharp boundary there, or that no observations were left out of the sample. The study design and data source matter when considering those possibilities.

Worked Example: Two Neighborhood Groups of Shuttle Times

A fictional community program records distance from a shuttle stop, \(x\), in kilometers, and ride time, \(y\), in minutes, for visitors from two neighborhoods. The invented data are below.

VisitorNeighborhoodDistance, \(x\) (km)Ride time, \(y\) (minutes)
ANorth142
BNorth248
CNorth355
DSouth764
ESouth870
FSouth974

Locate the clusters. The North visitors form a group at distances from 1 to 3 km and ride times from 42 to 55 minutes. The South visitors form another group at distances from 7 to 9 km and times from 64 to 74 minutes.

Describe the gaps. There are no observed distances between 3 and 7 km, so the groups are separated horizontally. The response values do not overlap: North visitors have times from 42 to 55 minutes and South visitors from 64 to 74 minutes, so there is also a gap in \(y\) between 55 and 64 minutes.

Consider what the pattern might indicate. The neighborhood labels match the two visible clusters. This suggests that neighborhood may be relevant to understanding the arrangement of these observations. The plot does not establish why the neighborhoods differ in distance or ride time, nor does it show that every visitor in either neighborhood has a value within the listed range.

Conclusion. The scatterplot has two clusters, one for each labeled neighborhood, separated in both distance and ride time. Describe those visible differences in context, but do not claim that neighborhood caused the differences based on this plot alone.

A Gap Along One Variable Can Still Separate Groups

Not every gap separates groups in both directions. Two clusters might have a clear horizontal gap in \(x\) while their \(y\)-values overlap, or a vertical gap in \(y\) while their \(x\)-values overlap. Being specific about the direction of the separation helps communicate what the scatterplot actually shows.

A horizontal gap means a range of explanatory-variable values has few or no observations. A vertical gap means a range of response-variable values has few or no observations. If the empty region is diagonal, describe the locations of the clouds rather than trying to force the gap into only one axis. In each case, keep the paired observations in view: a cluster is a region of the \(x\)-\(y\) plane, not just a group of values on one axis considered alone.

Worked Example: A Gap Between Two Practice-Time Groups

A fictional coach records practice time, \(x\), in minutes, and a skill score, \(y\), for two groups of athletes. Each point represents one athlete. The invented observations are shown below.

AthleteGroupPractice time, \(x\) (minutes)Skill score, \(y\)
AEarly session1018
BEarly session1220
CEarly session1421
DEarly session1623
ELate session3022
FLate session3224
GLate session3425
HLate session3627

Describe the arrangement. One cluster lies at practice times from 10 to 16 minutes, and another lies from 30 to 36 minutes. There are no observed practice times from 16 to 30 minutes, so there is a horizontal gap between the clusters.

Check whether there is also a vertical gap. The early-session skill scores range from 18 to 23, while the late-session scores range from 22 to 27. These ranges overlap from 22 to 23, so the groups are not separated by a gap in \(y\). This is an example of clusters separated in \(x\) without complete separation in \(y\).

Use the available labels cautiously. The session labels correspond to the two practice-time clusters. The plot supports saying that the recorded early-session athletes practiced for shorter times than the late-session athletes. It does not tell us why the sessions differ, or whether practice time alone explains the scores.

Conclusion. Describe the clear gap in practice time and the overlap in skill scores. Calling the groups separated in both variables would be inaccurate because some response values occur in both groups.

Clusters Can Reveal Structure, Not Just Separation

Sometimes the main clue is not a wide empty space but a change in how points are concentrated. A scatterplot might show a dense group in one region and a smaller, looser group nearby. It may also show multiple clusters with different shapes or directions. Look for these features before summarizing the display as one simple cloud.

Group labels can help you describe this structure, but they do not automatically explain it. If points are labeled by location, time, or category, check whether those labels line up with the visible clusters. If no label is available, describe the clusters by location in the plot and say that the reason for them is unknown from the graph alone. Avoid inventing a subgroup explanation just because it seems plausible.

Worked Example: Two Sampling Locations Along a River

A fictional ecology class records distance from a river mouth, \(x\), in kilometers, and the number of a certain type of plant in a sampling frame, \(y\), at eight sites. The locations are known to be either near the mouth or upstream. The values are invented for this example.

SiteLocation labelDistance, \(x\) (km)Plant count, \(y\)
ANear mouth28
BNear mouth310
CNear mouth49
DNear mouth512
EUpstream1615
FUpstream1718
GUpstream1816
HUpstream1920

Identify clusters and the gap. The near-mouth sites form a cluster from 2 to 5 km, with plant counts from 8 to 12. The upstream sites form another cluster from 16 to 19 km, with counts from 15 to 20. No sampled site lies between 5 and 16 km, leaving a horizontal gap between the clusters.

Connect the pattern to the labels. The labels indicate that the clusters correspond to the two sampled location types. A careful description is that the observed near-mouth sites are closer to the river mouth and have lower plant counts than the observed upstream sites in this sample.

Keep the limits clear. The plot does not show what plant counts would be at unsampled distances, and it does not establish why the sampled sites differ. The gap could reflect where the class chose to sample rather than a natural absence of sites or plants throughout that interval.

Conclusion. The graph shows two location-associated clusters with an unsampled distance range between them. The cluster labels provide useful context, but they do not prove a cause or describe every possible site along the river.

Common Mistakes and AP Exam Tips

  • Calling a gap proof that values cannot occur: A gap means few or no observations appear there in the displayed data. It does not establish that the values are impossible in the population.
  • Describing separation on the wrong axis: Check the ranges of both variables. Groups can be separated in \(x\) but overlap in \(y\), or the reverse. State exactly what the plot shows.
  • Assuming a cluster has a known cause: Labels may suggest a possible explanation, but a scatterplot alone does not establish why the observations form groups.
  • Treating a cluster as an outlier: A group of points concentrated together is not the same as one point far from a pattern. Use the distinction developed in “Spotting Outliers in Bivariate Data.”
  • Calling every sparse region a gap between groups: A small sample can look uneven by chance. Describe how many visible groups there are only when the plot supports that description.
  • Ignoring labels or study design: If observations are marked by group, setting, or time, use that information to investigate the plot. If such information is absent, do not guess it.
  • Using vague language: “There are two groups” is less informative than naming their approximate ranges in \(x\) and \(y\), and identifying any overlap or gap.

For a full-credit AP response, describe the location of each cluster using the variables, identify which variable or variables show a gap, and connect the pattern to context only as far as the evidence allows. For example: “The two neighborhood groups form separate clusters: North visitors have ride times from 42 to 55 minutes and South visitors from 64 to 74 minutes, leaving a gap in observed ride times. The plot shows an association between the labels and the clusters, but it does not establish why the groups differ.” That answer is specific and avoids treating a visual pattern as proof of a cause.

Key takeaway: Clusters are concentrations of paired observations, and gaps are sparse or empty regions between parts of a scatterplot. Describe where they occur in both variables, use labels and context carefully, and remember that the plot alone cannot explain why the structure exists.

Check Your Understanding

For each situation, describe the cluster or gap and explain what can—and cannot—be concluded from the scatterplot.

  1. A scatterplot has one cluster at \(x\)-values from 2 to 5 and a second at \(x\)-values from 12 to 16. Their \(y\)-values overlap. Which variable shows a gap, and which does not?
  2. Two labeled groups form clusters with no observations between them in either \(x\) or \(y\). What should you describe before suggesting what the labels might indicate?
  3. A point lies alone in a gap between two dense clusters. Why is it not automatically an error or an outlier?
  4. A sample has no points in a wide range of \(x\)-values. Does that establish that the population cannot have those values? Explain.
  5. A scatterplot shows two clusters, but the observations have no group labels. What can you say, and what explanation should you avoid claiming as fact?