Look Beyond the Overall Shape
In Reading Graphs for Shape, Gaps, and Outliers, you learned to make a feature inventory before writing about a quantitative distribution. This tutorial takes a closer look at three features that deserve careful attention: gaps, clusters, and potential outliers. These features can reveal where observations are concentrated and which observations stand apart, but they do not explain by themselves why the data look that way.
A gap is an interval with no observations. A cluster is a region where observations are concentrated relative to nearby regions. A potential outlier is an observation that stands apart from the rest of the data. “Potential” matters: a point that looks unusual in a display is a reason to investigate and describe it, not an automatic reason to delete it.
Use a display that shows the detail needed for the question. A dotplot or stemplot preserves individual observations, so isolated values and empty stretches can be easy to see. A histogram is useful for the overall pattern across intervals, but its appearance depends partly on the bin width and placement, as discussed in Choosing Class Width and Number of Bins. A boxplot can flag a possible outlier under a rule, but, as you learned in What a Boxplot Cannot Show, it does not reveal every gap or cluster.
A feature can also be unusual without being an outlier. For instance, a distribution may have two clusters separated by a gap. The clusters are important features even if neither one is an isolated point. Conversely, an observation may sit apart from the rest without the distribution having a clear gap or a second cluster.
A Practical Way to Identify and Report Features
Before reporting a gap, cluster, or potential outlier, check the display and the data’s context. Confirm what one observation represents, the units, and the scale. If a graph appears to show an isolated point, look at the original values if they are available. A single histogram bar with a small count identifies a sparsely populated interval, not necessarily the exact location of every observation in that interval.
Identify where most observations lie before deciding that a point or region stands apart.
Note empty intervals, concentrations separated by lower-count regions, and observations far from the main body.
Use the display’s scale and, when available, the individual data values. The 1.5 IQR rule from Applying the 1.5 IQR Rule for Outliers can flag values for further attention.
Name the group and variable, describe the feature and its location, and use cautious wording such as “possible outlier.” Do not invent a cause or discard a value just because it is unusual.
A clear report distinguishes observation from interpretation. “The histogram has no observations from 15 to 20 minutes” describes the data shown. “No wait can ever be between 15 and 20 minutes” goes beyond the data. Similarly, a gap between two clusters may suggest that the observations come from different subgroups, but the graph alone does not establish what those subgroups are or why they differ.
Reading a Gap and a Sparse Tail
Worked Example: Time Spent at a Community Workshop
A fictional community program records how many minutes 26 participants spend at a workshop. Its histogram uses intervals of equal width and shows these counts:
| Minutes | Number of participants |
|---|---|
| 0 to less than 5 | 8 |
| 5 to less than 10 | 10 |
| 10 to less than 15 | 7 |
| 15 to less than 20 | 0 |
| 20 to less than 25 | 1 |
Plan. Check that the bin counts account for the whole group, then identify any empty interval and any sparsely populated region. Since the histogram groups observations into intervals, report the lone observation by its interval rather than claiming to know its exact value.
Do. The counts add to \(8+10+7+0+1=26\), matching the 26 participants. There are no observations from 15 to less than 20 minutes. There is one observation from 20 to less than 25 minutes, separated from the main concentration by the empty interval. Most participants are in the first three intervals, which together contain \(8+10+7=25\) participants.
Conclude. For the 26 workshop participants, the histogram shows a gap from 15 to less than 20 minutes and one participant in the 20-to-less-than-25-minute interval, beyond the main concentration below 15 minutes. That observation may be unusual, but the histogram does not give its exact value or explain why that participant stayed longer.
The counts support a specific description: 25 participants spent less than 15 minutes, none spent 15 to less than 20 minutes, and one spent 20 to less than 25 minutes. They do not justify saying that the one participant’s time was exactly 20 minutes or exactly 25 minutes. Interval boundaries and the graph’s scale matter when describing a histogram.
Using the 1.5 IQR Rule as a Flag
A graph can make a value look separated from the rest, but the 1.5 IQR rule provides a numerical way to flag possible outliers. As explained in Applying the 1.5 IQR Rule for Outliers, values strictly below the lower fence or strictly above the upper fence are flagged. A flagged value still needs to be interpreted in context; the rule does not determine whether it is an error or should be removed.
Worked Example: A Long Wait at a Clinic
A fictional clinic records 12 patients’ waiting times, in minutes, on one afternoon. In order, the times are \(18, 19, 20, 20, 21, 22, 23, 24, 24, 25, 27,\) and \(61\). Use the median-of-halves method for quartiles described in Five-Number Summary and Boxplot Construction.
Plan. Find \(Q_1\) and \(Q_3\), calculate the IQR and fences, and check whether any observation lies strictly outside them. Then report what the rule flags without assuming that the recorded time is wrong.
Do. The lower six observations have middle values 20 and 20, so \(Q_1=(20+20)/2=20\) minutes. The upper six have middle values 24 and 25, so \(Q_3=(24+25)/2=24.5\) minutes. Therefore,
The lower fence is \(20-1.5(4.5)=20-6.75=13.25\) minutes. The upper fence is \(24.5+1.5(4.5)=24.5+6.75=31.25\) minutes. The smallest time, 18 minutes, is above the lower fence. The time of 61 minutes is greater than the upper fence of 31.25 minutes, so it is flagged.
Conclude. For these 12 patients, the 61-minute wait is a potential outlier by the 1.5 IQR rule. It is much longer than the other recorded waits, which range from 18 to 27 minutes. The rule identifies it for attention, but the clinic would need to check the record and context before deciding whether it is an error or a valid long wait.
The phrase “potential outlier by the 1.5 IQR rule” is more precise than simply calling a value “wrong.” A long wait could be correctly recorded and important to the clinic. If a value is checked and confirmed, include it when the goal is to describe the actual observations. If a recording error is discovered, explain how it was handled rather than silently removing it.
Clusters and Gaps May Point to Questions
Worked Example: Student Commute Times
A fictional survey records commute times, in minutes, for 18 students. The ordered values are \(4, 5, 5, 6, 6, 7, 7, 8, 8, 25, 26, 27, 28, 29, 30, 31, 32,\) and \(33\).
Plan. Examine where the observations are concentrated and where there are few or no observations. Describe the pattern in minutes and consider what question it might raise, without claiming that the data identify a cause.
Do. Nine commute times are between 4 and 8 minutes, and the other nine are between 25 and 33 minutes. There are no observations from 9 through 24 minutes. The values form two concentrations separated by a wide gap.
Conclude. Among the 18 students surveyed, commute times cluster around 4 to 8 minutes and 25 to 33 minutes, with no times from 9 through 24 minutes. This pattern raises a question about whether the students may represent different kinds of commutes, but the times alone do not show what explains the two clusters.
As you learned in Unimodal, Bimodal, and Multimodal Distributions, separate concentrations can be evidence of multiple peaks. Here, the ordered values make the two concentrations and the intervening gap clear. Still, a pattern that suggests subgroups is a prompt for further investigation, not proof that the groups differ in a particular way. Additional information about the students or how the data were collected would be needed to explore that idea.
Common Mistakes and AP Exam Tips
- Calling every extreme value an error. A value far from the rest may be a valid observation. Say “potential outlier” or “flagged by the 1.5 IQR rule,” then explain what should be checked.
- Overstating what a gap means. A gap means no observed values fall in that interval in this data set. It does not prove the variable cannot take values in that interval.
- Claiming a cause from a cluster. Clusters may suggest subgroups, but a graph alone does not identify those subgroups or explain why the pattern occurred.
- Reporting more precision than the display provides. A histogram bar identifies an interval, not the exact values within it. Give an exact observation only when the data or display supports it.
- Ignoring the scale or binning. Check the axis labels and intervals. A histogram’s apparent gaps and concentrations can look different with different bin choices, so describe the display accurately.
- Leaving out context. A strong description names the group, variable, units, and location of the feature. “One value is unusual” is less informative than reporting which value or interval stands apart and how it compares with the rest.
For full-credit communication, connect evidence to a cautious conclusion: “The histogram of workshop times shows no participants from 15 to less than 20 minutes and one from 20 to less than 25 minutes; that observation is separated from the main concentration, but the grouped graph does not reveal its exact value.” This identifies the feature, uses the units and intervals, and avoids an unsupported explanation.
Check Your Understanding
Use the display or values described in each question, and keep your conclusions tied to the evidence.
- A histogram has no observations from 30 to 40 minutes. What does that gap show, and what does it not establish?
- A dotplot has a main concentration between 6 and 12 and one observation at 31. Give a cautious way to describe the value at 31.
- A data set has two clusters separated by a gap. What can the graph suggest, and what can it not establish about the reason for the pattern?
- Under the 1.5 IQR rule, an observation lies above the upper fence. Does that alone mean it should be removed? Explain.
- A histogram has one observation in a bin from 50 to less than 60. Why should you avoid reporting its exact value as 50?