Spread Needs Context, Too
In Describing Center in Context, you learned to identify a mean or median, name the group and variable, and report the value with units. A center alone does not tell the whole story: two groups can have the same center but very different amounts of variability. A spread statement adds information about how much the observations differ from one another.
The range, interquartile range (IQR), and standard deviation are three measures of spread. They do not all describe variability in the same way. The range uses only the smallest and largest values, the IQR describes the middle half of the data, and standard deviation describes how far observations typically are from the mean. Each is reported in the original units of the variable.
As in The SOCS Framework for Describing Distributions, spread is one part of a description, alongside shape, unusual features, and center. A complete description keeps those features connected to the context. For example, “The IQR is 6 minutes” is more informative when the reader knows it describes the middle 50% of the recorded waiting times.
Range: The Full Span of the Observations
The range is the difference between the largest and smallest observations. It describes the full numerical span from one endpoint of the data to the other. Because it depends on only two observations, an unusually small or large value can change the range substantially.
Worked Example: Bike Commute Times
A fictional group records the times for six bike commutes to school: \(12, 14, 15, 16, 18,\) and \(25\) minutes. Interpret the range in context.
Plan. Find the largest and smallest recorded commute times, subtract the smaller value from the larger one, and explain what that difference represents.
Do. The minimum is 12 minutes, and the maximum is 25 minutes. Therefore,
A check is that the two endpoints are 13 minutes apart: counting upward from 12 to 25 gives a difference of 13, not 25.
Conclude. The six recorded bike commute times span 13 minutes, from the shortest time of 12 minutes to the longest time of 25 minutes. This does not mean each commute differed from a typical commute by 13 minutes; it describes only the distance between the two endpoints.
Range is quick to calculate and easy to interpret, but it gives no information about how the observations are distributed between the endpoints. A range of 13 minutes could come from values concentrated near the middle with one extreme observation, or from values spread fairly evenly across the interval. A dotplot or histogram can help reveal those differences.
IQR: Spread in the Middle Half
The interquartile range, or IQR, is \(Q_3-Q_1\). As you learned in Five-Number Summary and Boxplot Construction, \(Q_1\) and \(Q_3\) mark the first and third quartiles. The IQR measures the width of the interval containing the middle 50% of the observations. It does not describe the full span of the data.
An IQR of 7 points means the third quartile is 7 points above the first quartile. It does not mean every observation is within 7 points of the median, or that the middle half is centered exactly at the median. The quartiles mark the ends of the middle-half interval; the distribution within that interval may not be perfectly balanced.
Worked Example: Comparing Quiz-Score Variability
Two fictional groups each have eight recorded quiz scores. Group A’s ordered scores are \(18, 19, 20, 21, 22, 23, 24,\) and \(25\). Group B’s ordered scores are \(14, 18, 18, 21, 22, 24, 26,\) and \(30\). Compare their IQRs and interpret the difference.
Plan. Use the same quartile method for both groups: split each ordered list into a lower half and an upper half. Find \(Q_1\) as the median of the lower half and \(Q_3\) as the median of the upper half. Then subtract \(Q_1\) from \(Q_3\). With eight observations, each half has four values.
Do. For Group A, the lower half is \(18, 19, 20, 21\), so \(Q_1=(19+20)/2=19.5\) points. The upper half is \(22, 23, 24, 25\), so \(Q_3=(23+24)/2=23.5\) points. Thus,
For Group B, the lower half is \(14, 18, 18, 21\), so \(Q_1=(18+18)/2=18\) points. The upper half is \(22, 24, 26, 30\), so \(Q_3=(24+26)/2=25\) points. Therefore,
The subtraction can be checked by counting the distance between the quartile values: Group A’s quartiles are 4 points apart, and Group B’s are 7 points apart.
Conclude. The middle 50% of Group B’s recorded quiz scores spans 7 points, compared with 4 points for Group A. The scores in Group B therefore have greater spread in their middle halves. Both groups have a median of 21.5 points, so this comparison shows why reporting center without spread could leave out an important difference.
The IQR is less affected by an extreme observation than the range because the IQR is based on quartiles rather than the minimum and maximum. That makes it useful when describing the spread of a distribution with skew or unusual values. In Applying the 1.5 IQR Rule for Outliers, you also used the IQR to calculate fences; here, the goal is to interpret it as a measure of spread.
Standard Deviation: Typical Distance from the Mean
Standard deviation summarizes the typical distance of observations from their mean. It uses every observation, so it can describe overall variation more fully than the range. But unusually distant values can have a strong effect on it. Like the range and IQR, standard deviation is expressed in the same units as the data.
For a sample of \(n\) observations, the calculation of \(s\) uses the squared differences between each observation and the sample mean \(\bar{x}\), then takes a square root. For this tutorial, the key is what the resulting value means: a standard deviation of 3 hours describes observations that typically lie about 3 hours from the mean, not an exact distance for every observation.
Worked Example: Study Time for a Small Sample
A fictional sample of five students reports studying \(4, 6, 8, 10,\) and \(12\) hours during one weekend. Find the sample standard deviation and interpret it in context.
Plan. First find the sample mean. Then find each observation’s difference from the mean, square those differences, and use the sample standard deviation formula. Interpret the result as a typical distance from the mean, in hours.
Do. The total study time is \(4+6+8+10+12=40\) hours, so the sample mean is \(40/5=8\) hours. The deviations from 8 hours are \(-4,-2,0,2,\) and \(4\) hours; their squares are \(16,4,0,4,\) and \(16\) square hours. Their sum is 40 square hours. Therefore,
A calculation check uses the equivalent sum-of-squares form: the sum of the squared observations is \(4^2+6^2+8^2+10^2+12^2=360\), and the squared total divided by the sample size is \(40^2/5=320\). Their difference is \(360-320=40\); dividing by \(n-1=4\) gives 10, and \(\sqrt{10}\approx3.16\).
Conclude. The sample standard deviation of weekend study time for the five students is about 3.16 hours. Their study times typically differ from the sample mean of 8 hours by about 3.16 hours. This is a description of typical distance, not a claim that each student’s time is exactly 3.16 hours from 8.
The square root in the calculation returns the result to the original units. Squared differences are measured in square hours, but standard deviation is measured in hours. A standard deviation of 3.16 hours is therefore directly interpretable alongside the mean of 8 hours.
Choosing and Communicating a Spread Measure
Different measures answer different questions. Use the range when the full distance between the extremes is relevant. Use the IQR to focus on the width of the middle half, especially when unusual values or skew make the endpoints unrepresentative of most observations. Use standard deviation to summarize typical distance from the mean, particularly when the mean is a useful center and you want a measure that uses all the observations.
| Measure | What it summarizes | How unusual values can affect it |
|---|---|---|
| Range | Distance from the minimum to the maximum | Can change greatly if either endpoint is unusual |
| IQR | Width of the middle 50% of observations | Usually less affected by extreme values than the range |
| Standard deviation | Typical distance from the mean | Can be influenced by unusually distant observations |
The measurements have the same units as the variable, but a numerical value for spread is not automatically “large” or “small.” Interpret it alongside the variable, its context, and—when making a comparison—the spread in the other group. Comparisons are clearest when the groups measure the same variable in the same units. Do not compare, for example, a spread measured in minutes directly with one measured in kilograms.
Common Mistakes and AP Exam Tips
- Reporting a spread number without naming what it describes. “The IQR is 7” should identify the group, variable, and unit. Say, “The IQR of Group B’s quiz scores is 7 points.”
- Calling the IQR the distance of each observation from the median. The IQR is the width of the middle-half interval from \(Q_1\) to \(Q_3\); it is not a maximum distance from the median.
- Interpreting range as typical variation. Range describes the gap between the two extreme observations. It does not tell how far observations typically are from one another or from the center.
- Describing standard deviation as an exact distance for every value. Use wording such as “observations typically differ from the mean by about 3.16 hours.” Do not claim every observation is that far from the mean.
- Leaving out units. Report a range of 13 minutes, an IQR of 7 points, or a standard deviation of 3.16 hours—not just 13, 7, or 3.16.
- Making an unsupported population claim. A statistic calculated from recorded observations describes those data. Do not claim that it describes all people in a wider population unless the data collection supports that conclusion.
- Choosing a measure without considering unusual values. Range and standard deviation can respond strongly to extreme observations, while the IQR focuses on the middle half. Use the distribution’s features to help select a useful summary.
For full-credit communication, state what the measure summarizes and keep the statement in context. For example: “The sample standard deviation of weekend study time for the five students is about 3.16 hours, so their study times typically differ from the sample mean of 8 hours by about 3.16 hours.” This identifies the measure, group, variable, units, and meaning without implying every value has the same distance from the mean.
Check Your Understanding
For each question, explain what the spread measure says about the observations and include the appropriate units.
- A set of recorded water temperatures has a minimum of \(11^\circ\text{C}\) and a maximum of \(19^\circ\text{C}\). What is the range, and what does it represent?
- The first and third quartiles of a distribution of package masses are 2.4 kilograms and 3.1 kilograms. Find the IQR and interpret it.
- A group of runners has a mean training distance of 8 kilometers and a standard deviation of 1.5 kilometers. Write a contextual interpretation of the standard deviation.
- Which measure of spread focuses on the middle 50% of observations? Would an unusually large maximum usually affect it as much as the range?
- Why is “The range is 12” incomplete as a spread statement, even if the subtraction was done correctly?