Let Shape Guide Your Choice of Spread
In Choosing Mean or Median to Describe Center, you learned to use distribution shape and unusual values to choose a measure of center. The same features guide your choice of spread. As a general pairing, use the interquartile range (IQR) with the median for a skewed distribution or one with outliers. Use the standard deviation with the mean for a roughly symmetric distribution without strong outliers.
These pairings make sense because the measures respond differently to extreme values. The IQR describes the width of the middle half of the data and is resistant to a few extreme observations. Standard deviation describes a typical distance from the mean and is non-resistant: an extreme observation can substantially affect it. Neither measure is automatically wrong in a particular data set; the goal is to choose the measure that best summarizes the distribution for the question at hand.
As covered in Describing Spread in Context, the IQR is \(Q_3-Q_1\), and sample standard deviation is denoted by \(s\). The IQR has the same units as the observations. The standard deviation also has the same units as the observations; it describes a typical distance from the mean, rather than the width of a specified portion of the data.
A Practical Decision Process
Before choosing a measure, look at an appropriate graph, such as a dotplot or histogram. Use the SOCS approach from The SOCS Framework for Describing Distributions: shape and outliers are particularly relevant to this decision. A five-number summary can help with the IQR, but a graph can reveal skew or isolated values that affect how useful a measure will be.
Decide whether it is approximately symmetric or skewed, and check whether any observations stand apart as possible outliers.
Use the IQR as a resistant summary for skew or outliers. Use standard deviation when the distribution is approximately symmetric and has no strong outliers.
Report the measure with units. Explain what it describes: the width of the middle 50% for IQR, or a typical distance from the mean for standard deviation.
“Resistant” means relatively unaffected by a few extreme values, not completely unchanged by every alteration to the data. An extreme value can change quartiles in some data sets, but the IQR generally responds less to that value than the standard deviation does. Standard deviation uses distances from the mean, and those distances are squared in its calculation; a very distant observation can therefore have a substantial effect.
Skew and outliers are related clues, but they are not the same thing. A distribution can be skewed without one observation standing far apart, and an observation can be unusual even when the overall pattern looks fairly balanced. Consider both the overall shape and any unusual observations before settling on a summary.
Worked Examples
Worked Example: Symmetric Battery Lifetimes
A fictional sample of eight rechargeable batteries lasts \(12, 14, 16, 18, 20, 22, 24,\) and \(26\) hours. A dotplot shows a roughly symmetric pattern and no strong outliers. Which measure of spread fits, and what does it say about the lifetimes?
State. The distribution is approximately symmetric and has no strong outliers, so standard deviation is a suitable measure of spread.
Plan. Calculate the sample standard deviation from the mean and the squared deviations. Also find the IQR for comparison, using the quartile method described in Five-Number Summary and Boxplot Construction. Then interpret the standard deviation in hours.
Do. The mean lifetime is \((12+14+16+18+20+22+24+26)/8=152/8=19\) hours. The squared deviations from 19 are \(49, 25, 9, 1, 1, 9, 25,\) and \(49\), which sum to \(168\). Thus:
The sum of squared deviations can be checked by pairing equal squares: \(49+49=98\), \(25+25=50\), and \(9+9+1+1=20\), for a total of \(168\). For the IQR, the lower half is \(12,14,16,18\), so \(Q_1=(14+16)/2=15\). The upper half is \(20,22,24,26\), so \(Q_3=(22+24)/2=23\). Therefore, \(\text{IQR}=23-15=8\) hours.
Conclude. The sample standard deviation of battery lifetime is about \(4.90\) hours. In this roughly symmetric sample without strong outliers, it is reasonable to describe lifetimes as typically about \(4.90\) hours from the mean of 19 hours. The IQR is also a valid calculation: the middle half of the observed lifetimes spans 8 hours.
Worked Example: A Long Delivery Time
A fictional delivery service records nine delivery times, in minutes: \(3, 4, 4, 5, 5, 6, 7, 9,\) and \(20\). The graph shows a cluster of shorter deliveries and a long right tail. Choose a useful measure of spread and explain why.
State. The distribution is right-skewed, with one unusually long delivery. The IQR is a useful, resistant summary of spread.
Plan. Find \(Q_1\) and \(Q_3\) from the ordered values and calculate the IQR. Calculate \(s\) as well to illustrate how the long delivery affects a non-resistant measure. Use the 1.5 IQR rule from Applying the 1.5 IQR Rule for Outliers to check whether 20 minutes is flagged.
Do. The median is the fifth value, 5 minutes. The lower half, \(3,4,4,5\), has \(Q_1=(4+4)/2=4\) minutes. The upper half, \(6,7,9,20\), has \(Q_3=(7+9)/2=8\) minutes. Therefore:
The fences are \(Q_1-1.5(\text{IQR})=4-6=-2\) minutes and \(Q_3+1.5(\text{IQR})=8+6=14\) minutes. Since \(20>14\), the 20-minute delivery is flagged by the 1.5 IQR rule. This rule identifies a value to investigate; it does not prove that the observation is an error.
To calculate \(s\), the mean is \(63/9=7\) minutes. The squared deviations from 7 are \(16,9,9,4,4,1,0,4,\) and \(169\), summing to \(216\). Hence:
As a check, the values sum to \(3+4+4+5+5+6+7+9+20=63\), and the squared deviations sum to \(16+9+9+4+4+1+0+4+169=216\).
Conclude. The IQR of the nine delivery times is 4 minutes, so the middle half of the observed times spans 4 minutes. The standard deviation is about 5.20 minutes, but the long delivery contributes a large squared deviation and makes \(s\) less representative of the spread among the clustered shorter deliveries. The IQR is the more useful summary for this skewed distribution.
Worked Example: Comparing Two Balanced Groups
Two fictional groups test the battery life of a new device. Group A records \(31,33,35,37,39,41,43,\) and \(45\) hours. Group B records \(34,35,36,37,39,40,41,\) and \(42\) hours. Both dotplots are roughly symmetric, with no strong outliers. Which spread measure should be used to compare the groups, and which group has more variability?
State. Both distributions are approximately symmetric without strong outliers, so standard deviation is appropriate for comparing their variability.
Plan. Find the sample standard deviation for each group using the sample standard deviation formula. Check the IQRs too, then compare the standard deviations because the distribution shapes support that choice. The observations are battery lifetimes, so report spread in hours.
Do. Group A has mean \(38\) hours. Its deviations from 38 are \(-7,-5,-3,-1,1,3,5,\) and \(7\), whose squares sum to \(168\). Thus \(s_A=\sqrt{168/7}=\sqrt{24}\approx4.90\) hours. Group B also has mean \(38\) hours. Its squared deviations sum to \(60\), so \(s_B=\sqrt{60/7}\approx2.93\) hours. To check the total for Group B, the squared deviations are \(16,9,4,1,1,4,9,\) and \(16\), which sum to \(60\).
For Group A, \(Q_1=(33+35)/2=34\) hours and \(Q_3=(41+43)/2=42\) hours, giving an IQR of \(42-34=8\) hours. For Group B, \(Q_1=(35+36)/2=35.5\) hours and \(Q_3=(40+41)/2=40.5\) hours, giving an IQR of \(40.5-35.5=5\) hours.
Conclude. Group A’s sample standard deviation is about \(4.90\) hours, compared with about \(2.93\) hours for Group B. The measurements in Group A are more variable by this measure. Both IQRs support the same comparison: the middle half spans 8 hours in Group A and 5 hours in Group B. The numerical values of IQR and standard deviation describe different features, so compare like measures rather than treating an IQR of 8 hours as directly equivalent to an \(s\) of 8 hours.
Using Spread Measures Carefully
A measure of spread should be interpreted according to what it summarizes. An IQR of 4 minutes does not mean every delivery time is within 4 minutes of the median. It means the distance from the first quartile to the third quartile is 4 minutes. A standard deviation of 4.90 hours does not mean every battery lasts within 4.90 hours of the mean. It describes a typical distance from the mean, and some observations may be farther away.
When you compare two groups, inspect both distributions before comparing their spread. If both are roughly symmetric without strong outliers, comparing their standard deviations is sensible. If both are skewed or contain outliers, compare their IQRs. If the groups have different shapes, explain which measure is appropriate for each and be cautious about a direct comparison: the measures summarize spread in different ways.
These are guidelines for informative description, not rules that forbid calculating another measure. A question might ask specifically for a standard deviation or IQR; calculate what is requested, then explain any limitation its distribution shape creates. A skewed distribution does not make standard deviation mathematically invalid. It can make standard deviation a less helpful description of the spread of most observations.
Common Mistakes and AP Exam Tips
- Choosing a measure without inspecting shape. A full-credit response links the choice to evidence, such as “The distribution is right-skewed, so the resistant IQR is a useful summary of spread.”
- Calling standard deviation resistant. Standard deviation is non-resistant because extreme values can change it substantially. The IQR is generally resistant to a few extreme values.
- Describing the IQR as a distance from the median. The IQR is the width of the middle 50%, from \(Q_1\) to \(Q_3\). It is not the average distance from the median.
- Claiming standard deviation gives the exact distance of every observation from the mean. It describes a typical distance. Individual observations can be closer to or farther from the mean.
- Assuming a flagged outlier must be removed. The 1.5 IQR rule flags observations for attention; it does not establish that they are mistakes. Keep valid observations unless there is a reason, in context, to exclude them.
- Comparing unlike measures as if they were interchangeable. Both measures use the original units, but the IQR summarizes the middle half while standard deviation summarizes typical distance from the mean. Compare the same measure across groups when possible.
- Leaving out context and units. State what variable and group the measure describes, include units, and say what the number represents.
Check Your Understanding
For each situation, decide which spread measure is more useful and explain what it describes.
- A roughly symmetric distribution of water-bottle fill volumes has no strong outliers. Would you report the IQR or standard deviation as the main measure of spread? Explain.
- A set of household electricity bills is strongly right-skewed because a few bills are much larger than most. Which measure is likely to summarize the spread of the middle of the data more usefully?
- A sample has \(Q_1=12\) minutes and \(Q_3=20\) minutes. Find the IQR and state what it means in context if the variable is time spent waiting for a bus.
- One very large value is added to a data set. Which measure of spread is generally more affected, the IQR or the standard deviation? Why?
- Two groups are both roughly symmetric and have no strong outliers. Group A has \(s=3.2\) seconds and Group B has \(s=5.1\) seconds. Which group has greater variability by standard deviation, and what units should appear in your answer?