Scan First, Describe Second
A graph can show several features at once. If you start writing immediately, it is easy to overlook a gap, call an ordinary value an outlier, or describe the shape without referring to what the graph actually shows. A useful habit is to scan the display and list its visible features first. Then turn that inventory into a short description in context.
In Reading Values and Counts From a Dotplot, Reading Counts and Percents From a Histogram, and Reading a Boxplot's Quartiles and Spread, you learned how to read values or summaries from each display. This tutorial focuses on what to look for across quantitative graphs: the overall shape, concentrations, gaps, and possible outliers. These features are not equally visible in every type of graph. A dotplot shows individual observations; a histogram groups observations into intervals; a boxplot summarizes positions and spread.
Start by identifying the variable, its units, the group represented, and the type of display. Then scan from left to right. Ask where observations are concentrated, whether the distribution has one or more peaks, whether it stretches farther in one direction, whether there are empty intervals, and whether any values are separated from the rest. Record only features the display supports.
Name the quantitative variable, its units, and the group whose data are shown.
Look for one main concentration or several, and decide whether the pattern is roughly symmetric or has a longer tail in one direction.
Note clusters, peaks, sparse regions, empty intervals, and values that appear separated from the main body.
Use the feature list to describe the distribution in context. Include visible values when they help, and qualify uncertain claims.
Shape, Peaks, and Tails
The shape is the distribution's broad visual pattern. A distribution may be roughly symmetric, meaning its two sides have broadly similar shape and spread around a central region. It may be skewed, with a longer tail on one side. A distribution with a longer tail toward larger values is described as skewed to the right; a longer tail toward smaller values is skewed to the left. The direction of skew is named for the tail, not for where most of the observations are concentrated.
A peak is a local high point or concentration in a display. A distribution with one main peak is often called unimodal; one with two distinct peaks is often called bimodal. A cluster is a region where observations are concentrated. A cluster can cover a range of values rather than form one sharp peak. Do not count every small rise and dip as a separate peak: sampling variation and the choice of histogram bins can create minor bumps.
The graph type affects how confidently you can describe shape. Individual values in a dotplot may show the pattern clearly, while histogram bin choices can smooth over or exaggerate local bumps. A boxplot does not display peaks or tails in enough detail to identify modality or skew reliably. As explained in What a Boxplot Cannot Show, use a histogram or dotplot when you need to inspect those features.
Gaps and Possible Outliers
A gap is an interval containing no observations. A dotplot can show a run of values with no dots between two groups. A histogram can show an empty bin or a span of empty bins. A low bar is not necessarily a gap: if the bar has positive height, observations fell in that interval. Also, the apparent location of a histogram gap depends on its bins, so describe the interval shown rather than claiming that the graph identifies every exact value that is absent.
A value far from the rest of a distribution may look like a possible outlier. “Possible” matters: a visual impression is a prompt to inspect the value, not proof that it is an error or should be removed. A dotplot may show an isolated point directly. A modified boxplot may mark values beyond its whiskers as individual points. A basic boxplot, however, may not show individual observations beyond its whiskers. In Applying the 1.5 IQR Rule for Outliers, you learned a numerical rule for flagging potential outliers; use that rule when the task requires a formal check rather than relying on appearance alone.
Worked Example: Make an Inventory From a Dotplot
A fictional parks group records the duration, in minutes, of 21 short walking trips. The dotplot has the following number of dots above each value:
| Trip duration (minutes) | Number of dots |
|---|---|
| 5 | 1 |
| 10 | 2 |
| 15 | 4 |
| 20 | 5 |
| 25 | 4 |
| 30 | 3 |
| 35 | 1 |
| 60 | 1 |
Plan. Identify the main concentration, check the overall pattern and tails, and look for an empty span or an isolated value. Then write a description supported by the displayed values.
Do. The counts sum to \(1+2+4+5+4+3+1+1=21\), matching the stated number of trips. Most observations are between 10 and 35 minutes, with the greatest stack at 20 minutes. The stacks increase toward 20 and then generally decrease through 35, forming one main concentration. There are no dots from 36 through 59 minutes, and the dot at 60 is separated from the main group. The values near 5 and 60 make the pattern extend farther toward larger values; the isolated 60-minute trip gives the display a long high-value tail.
Feature inventory: one main concentration from about 10 to 35 minutes; highest stack at 20 minutes; an empty span between 35 and 60; one isolated value at 60; a longer extension toward larger durations.
Conclude. The walking-trip durations are concentrated between about 10 and 35 minutes, with the most trips at 20 minutes. The dotplot also shows no trips between 35 and 60 minutes and one isolated 60-minute trip, so the distribution extends farther toward larger durations. The graph alone does not establish why that trip took longer.
Read Histogram Features With the Bins in Mind
A histogram's shape, clusters, and gaps are visible at the scale of its intervals. A bar with a small positive height represents a bin with some observations, not an empty interval. If neighboring bins have different heights, that change can help locate peaks or sparse regions, but it does not reveal the exact positions of observations inside each bin.
Before describing a gap in a histogram, check the bin boundaries and whether the bin has zero frequency. The earlier tutorial Common Errors in Drawing Histograms explains why a genuine blank span should represent an empty interval, not unwanted spacing between bars. Also be cautious about declaring a distribution bimodal from a single dip: the apparent number of peaks can change when bin widths or starting boundaries change.
Worked Example: Distinguish a Gap From a Low Bar
A fictional garden project records the number of minutes 40 seedlings take to emerge after watering. A histogram uses equal-width bins:
| Time to emerge (minutes) | Frequency |
|---|---|
| [0, 2) | 8 |
| [2, 4) | 10 |
| [4, 6) | 2 |
| [6, 8) | 0 |
| [8, 10) | 1 |
| [10, 12) | 7 |
| [12, 14) | 9 |
| [14, 16) | 3 |
Plan. Check the total, identify the main concentrations, and distinguish the bin with zero frequency from the low but nonzero bars. Describe the visible pattern at the resolution of these two-minute bins.
Do. The frequencies sum to \(8+10+2+0+1+7+9+3=40\), matching the stated sample size. The larger bars are in \([0,4)\) and \([10,14)\), with a sparse region between these concentrations. The \([4,6)\) bin has frequency 2, so it is low but not empty. The \([6,8)\) bin has frequency 0, so it is the only empty bin. The \([8,10)\) bin has frequency 1, so there is an observation there; the entire interval from 4 to 10 minutes is therefore not empty.
Feature inventory: two broad concentrations, one near 0–4 minutes and another near 10–14 minutes; a sparse middle region; one empty bin from 6 to 8 minutes; no evidence in this display that every value from 4 to 10 minutes is absent.
Conclude. The histogram suggests two concentrations of emergence times, one from 0 to 4 minutes and another from 10 to 14 minutes, separated by a sparse middle region. Only the 6-to-8-minute bin is empty; the bars show observations in both the 4-to-6 and 8-to-10-minute bins. Because the data are grouped, the histogram does not show the exact values within any bin.
What a Boxplot Can—and Cannot—Add
A boxplot is useful for scanning the distribution's center and spread, and a modified boxplot can make flagged observations visible. Include the median, quartiles, and whisker endpoints in an inventory when they are relevant. But do not use a boxplot to claim that a distribution has a particular number of peaks or that a gap exists between two values: the box and whiskers do not preserve that detail.
An isolated point on a modified boxplot can be reported as a plotted potential outlier. If a question asks whether the value meets the 1.5 IQR rule, use the calculation from Applying the 1.5 IQR Rule for Outliers. A boxplot that does not plot individual outliers does not let you determine from its whisker alone whether there is a gap beyond the whisker or how many observations lie there.
Worked Example: Describe a Modified Boxplot Cautiously
A fictional modified boxplot shows travel distances, in kilometers, for a group of delivery routes. Its left whisker ends at 6, the left edge of the box is at \(Q_1=12\), the median is 16, the right edge is at \(Q_3=21\), and the right whisker ends at 29. One individual point is plotted at 48.
Plan. List the values and features a boxplot actually shows. Treat the plotted point as a possible outlier, and do not infer peaks or gaps that the display cannot reveal.
Do. The middle half of the displayed distances lies between 12 and 21 kilometers, and the median is 16 kilometers. The whiskers reach 6 and 29 kilometers, while the modified boxplot separately marks a route at 48 kilometers. To check that point using the previously learned 1.5 IQR rule, the IQR is \(21-12=9\) kilometers and the upper fence is \(21+1.5(9)=34.5\) kilometers. Since \(48>34.5\), the displayed route is above the upper fence.
Feature inventory: median 16 kilometers; middle 50% from 12 to 21 kilometers; whiskers from 6 to 29 kilometers; one plotted point at 48 kilometers, above the upper fence. The boxplot does not reveal how many peaks the distances have or whether there are gaps between the whisker and the point.
Conclude. The delivery-route distances have a median of 16 kilometers, with the middle half between 12 and 21 kilometers. The modified boxplot marks one 48-kilometer route beyond the upper fence as a potential high outlier. The boxplot alone does not show the distribution's detailed shape or the reason for that long route.
Common Mistakes and AP Exam Tips
- Writing a label instead of a description. “Skewed” is incomplete. State the direction of the longer tail and identify the variable in context.
- Calling the main cluster the tail. A right-skewed distribution has its longer tail toward larger values, even though most observations may be on the lower-value side.
- Calling a low bar a gap. A positive-height histogram bar means the interval contains observations. A gap requires an empty interval.
- Overstating what the display shows. A histogram groups values into bins, and a boxplot hides peaks and gaps. Describe what is visible at the display's resolution.
- Calling every distant observation an error. Say “possible outlier” or “appears isolated.” A graph does not explain why the value is unusual or whether it should be removed.
- Listing features without context. Name the measured variable and units, such as “trip durations in minutes,” rather than writing only “one cluster and one outlier.”
A strong response first gives an accurate feature inventory and then connects the features in a concise description. For full-credit communication, identify the variable and group, describe the overall shape, state where concentrations or gaps occur using the graph's scale, and qualify a visually unusual point. Avoid claiming that a graph proves a cause or supplies details it does not display.
Check Your Understanding
For each question, use a feature inventory before composing a description.
- A dotplot has most values between 12 and 20, with one point at 47 and no dots from 21 through 46. Name two features to include and give cautious wording for the point at 47.
- A histogram has a bar of height 1 from 8 to 10 minutes and a bar of height 0 from 10 to 12 minutes. Which interval is a gap, and why is the other interval not empty?
- A distribution has its main concentration near smaller values and a long extension toward larger values. How should its skew be named?
- A boxplot shows a long whisker and a plotted point beyond it. Name one conclusion the display supports and one claim about shape or gaps it does not support.
- Why should a description of histogram peaks mention that the graph groups observations into bins?