Small Wording Choices Can Change a Description
In Effect of Removing an Outlier, you compared summaries calculated with and without an unusual observation. Here the focus is on describing a distribution clearly. Even when you recognize a graph’s features, a vague sentence or a reversed skew label can make your description inaccurate or hard to interpret.
As in The SOCS Framework for Describing Distributions, a useful description considers shape, outliers, center, and spread. In Writing a Complete Description of a Distribution, you learned to name the group and variable and include units when reporting summaries. This tutorial concentrates on errors that can creep into those descriptions: saying too little, leaving out context, using “skewed to the right” incorrectly, and confusing the direction of a distribution’s tail.
Four Checks Before You Describe a Distribution
A clear description is not just a list of adjectives. Before writing, check what the graph displays and what you can reasonably say about it.
Say whose or what observations are shown and what was measured. “Students’ quiz scores” is more informative than “the data.”
Replace words such as “weird,” “high,” or “spread out” with supported descriptions of shape, a cluster, a gap, an isolated value, center, or spread.
Follow the less-concentrated end of the distribution along the number line. A tail toward larger values indicates right skew; a tail toward smaller values indicates left skew.
When reporting a center or spread, name the measure and give its value with units. When describing shape, connect it to the variable and group.
“Vague” does not mean that a statement is necessarily false. It means the reader cannot tell exactly what feature you noticed or what values it concerns. For example, “the scores are spread out” does not say whether you mean the full range, the middle half, or typical distance from the mean. As in Describing Spread in Context, name the measure if you are making a claim about spread.
Context matters for the same reason. “The median is 6” leaves the reader to guess what is being measured, which group the value describes, and whether 6 means points, minutes, or some other unit. A contextual statement might say, “For the nine students in this fictional group, the median quiz score is 6 points.” Shape statements also need context: “The distribution of delivery times is right-skewed” identifies what the pattern describes, while “it is skewed” does not.
Skew Is Named for the Tail
To check skew, imagine moving along the horizontal axis from smaller values to larger values. Find the main concentration of observations, then look for the thinner, more extended end of the distribution. The peak is where observations are concentrated; it is not the part that gives skew its name.
A common source of confusion is thinking that a distribution is right-skewed because most observations are on the right side of the graph. In fact, a right-skewed distribution often has most observations at smaller values and a few observations extending toward larger values. The reverse can occur for a left-skewed distribution. As described in Describing Shape: Symmetric, Skewed, Uniform, name skew from the tail, not simply from where the observations pile up.
A display can be drawn or presented in different ways, so first check its axes. In the usual histogram, dotplot, or stemplot, numerical values increase from left to right. The tail’s direction is still determined by whether it reaches toward larger or smaller values—not by where the tallest bars or densest stacks appear.
Worked Examples
Worked Example: The Peak Is Not the Skew Direction
A fictional repair desk records how long customers wait, in minutes. The dotplot is summarized below by the number of observations at each value.
| Wait time (minutes) | 2 | 3 | 4 | 5 | 6 | 9 | 14 | 22 |
|---|---|---|---|---|---|---|---|---|
| Number of customers | 4 | 6 | 7 | 5 | 3 | 1 | 1 | 1 |
A student says, “The distribution is left-skewed because most of the customers are on the left.” Is this description correct?
Solution. The largest concentration is between 2 and 6 minutes, and the most common wait is 4 minutes. Beyond that concentration, a few observations continue toward larger values: 9, 14, and 22 minutes. There is no comparable tail extending below 2 minutes. The thinner, more extended end therefore points toward larger values.
Conclusion. The distribution of customer wait times is right-skewed, not left-skewed. The student noticed that many observations lie at relatively small values, but used that fact to name the direction incorrectly. The direction comes from the tail toward larger wait times.
Worked Example: Replace a Vague Statement With Supported Details
In a fictional group of nine Grade 10 students, quiz scores out of 15 points are 4, 5, 5, 6, 6, 7, 8, 9, and 14. A student writes, “The scores are pretty high, and the data are spread out.” Identify what is missing and use the observations to make a more informative description.
Solution. “Pretty high” does not identify a reference point, and “spread out” does not name a measure of spread or indicate how the values are arranged. The statement also does not say which group or variable the observations describe, or give the units. We can check the ordered values for a center and spread summary. There are nine scores, so the fifth score is the median: 6 points.
Using the median-of-halves method from earlier in this course, the lower four scores are 4, 5, 5, and 6, so \(Q_1=(5+5)/2=5\) points. The upper four are 7, 8, 9, and 14, so \(Q_3=(8+9)/2=8.5\) points. Thus, the IQR is \(8.5-5=3.5\) points. The value of 14 is separated from the next-highest score, 9, in this list.
Conclusion. A clearer description is: “For these nine fictional Grade 10 students, quiz scores are concentrated from 4 to 9 points, with a median of 6 points and an IQR of 3.5 points. The score of 14 points is separated from the next-highest score.” This version names the group, variable, and units, and reports specific values rather than calling the scores “high” or “spread out.” It describes the visible separation without claiming, on that basis alone, why the score occurred.
Worked Example: A High Peak Can Have a Left Tail
A fictional teacher summarizes scores on a 20-point exit ticket that most students found accessible. The score axis runs from 0 to 20, and a dotplot has the following counts.
| Score (points) | 8 | 12 | 15 | 17 | 18 | 19 | 20 |
|---|---|---|---|---|---|---|---|
| Number of students | 1 | 1 | 2 | 4 | 6 | 7 | 5 |
A student says, “It is right-skewed because the peak is on the right.” Check the claim using the direction of the tail.
Solution. The greatest concentration is at higher scores: 17 through 20 points, with the peak at 19 points. The lower counts continue toward smaller values, including 15, 12, and 8 points. On the numerical axis, that less-concentrated end extends to the left of the main concentration. There is no similarly extended tail toward values above 20.
Conclusion. This distribution is left-skewed, not right-skewed. The student correctly noticed that the peak is on the right, but a peak’s position does not name the skew direction. Here the tail extends toward smaller scores.
Worked Example: Avoid Claiming More Than the Graph Shows
A fictional dotplot displays the number of minutes that 12 volunteers spend sorting donated books during one session. The values are 18, 19, 20, 20, 21, 21, 21, 22, 22, 23, 24, and 35 minutes. A student writes, “Everyone usually takes about 21 minutes, except for an error at 35.”
Solution. The list shows a concentration from 18 to 24 minutes and a separate larger value at 35 minutes. It does not show that everyone takes about 21 minutes: the observed values vary, and the data describe these 12 volunteers, not necessarily all volunteers or future sessions. Nor does the dotplot establish that 35 is an error. As in What Outliers Mean in Context, a value that stands apart should not be treated as a mistake without evidence about its source.
Conclusion. A supported description is: “For the 12 fictional volunteers in this session, sorting times are concentrated from 18 to 24 minutes, with one higher value of 35 minutes.” This sentence names the group, variable, and units and distinguishes what the display shows from an unsupported explanation. It does not say that every volunteer takes about 21 minutes or that 35 minutes is an error.
Common Mistakes and AP Exam Tips
- Using “the data” without naming the variable. A reader should not have to infer what was measured. Name the group and quantitative variable, such as “the distribution of book-sorting times for these volunteers.”
- Using vague adjectives in place of evidence. “High,” “low,” “wide,” or “weird” may express an impression but do not identify a feature clearly. Refer to observed values, a cluster, a gap, a tail, a center, or a named measure of spread.
- Naming skew from the busiest side. A peak on the right does not make a distribution right-skewed. Trace the tail along the numerical axis before naming skew.
- Thinking right-skew means most observations are large. A right tail points toward large values, but the main concentration can be at smaller values. Describe both the concentration and the tail when that distinction helps clarify the shape.
- Calling a flagged value an error. A graph can show that a value stands apart; it cannot, on its own, establish that the value was recorded incorrectly. Describe what is visible and avoid inventing a cause.
- Forgetting context or units. A median, IQR, range, or standard deviation needs the variable, group, and units. For full-credit communication, a reader should understand what the number describes without guessing.
- Making a broader claim than the observations support. If a graph describes a particular group, say so. Avoid changing “these students’ scores” into a statement about all students unless the evidence and question support that wider scope.
Check Your Understanding
Use the observations and context given in each question. Focus on what the display supports and, when asked about skew, the direction of the tail.
- A histogram of fictional commute times has most observations between 10 and 20 minutes, with a few observations extending toward 50 minutes. Which direction is the tail, and what is the skew description?
- A dotplot shows test scores clustered near 90 points, with a few much smaller scores extending toward 40. A student calls the distribution right-skewed because its peak is near the right side. Explain the error.
- Why is “the data are spread out” too vague? Name two details a more informative statement could include.
- A student reports, “The median is 12.” What context and measurement information could be missing from that statement?
- A dotplot shows one value far from the main cluster. What can you describe from the display, and what would you need to know before calling that value an error?