Tutorials › AP Statistics › Why Comparing Distributions Needs All Four Features

Comparing distributions · Tutorial 121 of 1000

Why Comparing Distributions Needs All Four Features

Practice making side-by-side comparisons that address shape, center, spread, and outliers for each group.

Beginner 9 min read

What You'll Learn

  • Organize a comparison so all four distribution features are addressed for both groups.
  • Compare centers using an appropriate measure and report the difference with units.
  • Compare spread without confusing the IQR, range, and standard deviation.
  • Describe how shape and unusual observations differ between groups.
  • Explain why matching centers do not necessarily mean two distributions are alike.

What Makes a Comparison Complete?

When two groups have measurements of the same quantitative variable, it is tempting to compare just their averages or medians. But center alone can hide important differences. One group might have values clustered tightly around its center, while another has a much wider spread or an unusual observation. Their shapes might differ, too.

In The SOCS Framework for Describing Distributions, you learned to describe Shape, Outliers, Center, and Spread. A comparison puts that framework to work twice: describe each group’s distribution, then explain how the features are alike or different. A useful comparison keeps the variable, units, and groups clear throughout.

Key idea: A valid comparison addresses shape, center, spread, and outliers in each group. It then states the important similarities and differences in context, using evidence from graphs or summaries.

“Group A has a larger median” is a comparison of center, not a complete comparison of distributions. A fuller response might also say whether the groups have similar shapes, which has greater spread, and whether either group has an unusual value. If no outliers are apparent, say so rather than leaving that feature unaddressed.

A Side-by-Side Comparison Method

Start by checking that both distributions describe the same variable in the same units. Then examine a graph, such as a dotplot, histogram, or boxplot, along with relevant summary statistics. As in Describing a Distribution With Real Numbers, use graph scales and summaries accurately; do not claim more detail than the display supports.

1
Describe shape for both groups.
Compare features such as symmetry, skew, peaks, clusters, or gaps. Use the graph to support the description.
2
Compare centers.
Choose a suitable measure, such as the median or mean, and report both values. State which is larger and by how much when that is useful.
3
Compare spread.
Use a suitable measure for each distribution, such as IQR or standard deviation. Keep the measure and units clear; do not compare unlike measures as if they were the same.
4
Check outliers and unusual features.
Report supported observations that stand apart, or state that neither group shows an apparent outlier. Do not assume an unusual value is an error.
5
Write the comparison in context.
Connect the features to the groups and variable. Make clear which group has the higher center, greater spread, or different shape.

The order of your sentences can vary. You might discuss each group in turn using SOCS, or compare the groups feature by feature. Either way, cover all four features for both groups. A table can help you organize what the evidence says before you write a paragraph.

Worked Example: Comparing Two Routes’ Travel Times

Worked Example: Comparing Two Routes’ Travel Times

A fictional student group recorded how long a trip to school took, in minutes, using either Route A or Route B. The observations for each route are listed in order:

Route A (minutes)Route B (minutes)
12, 13, 14, 14, 15, 15, 16, 16, 17, 189, 10, 11, 12, 13, 14, 16, 18, 22, 31

Compare the two distributions using shape, center, spread, and outliers.

State. The groups are students using Route A and Route B. The quantitative variable is trip time to school, measured in minutes. These observations describe the students in this fictional group.

Plan. We can compare the ordered lists directly. Use the median and IQR to describe center and spread because Route B appears right-skewed and includes a value that may be unusual. As in Finding Quartiles and the IQR, use the median-of-halves convention to find quartiles. Check a possible outlier with the 1.5 IQR rule, as discussed in Identifying Outliers and Unusual Features.

Do: shape and center. Route A’s times are fairly balanced around the middle, with no long tail in either direction apparent in this small list. Its median is the average of the fifth and sixth values:

$$ \text{Median}_A=\frac{15+15}{2}=15\text{ minutes} $$

Route B has a longer tail toward larger times: after several values from 9 to 18 minutes, the observations extend to 22 and 31 minutes. This indicates right skew. Its median is:

$$ \text{Median}_B=\frac{13+14}{2}=13.5\text{ minutes} $$

Route A’s median is \(15-13.5=1.5\) minutes greater. In this group, the typical trip time by median was 1.5 minutes longer on Route A than on Route B. That center comparison does not mean Route A’s times are generally more consistent.

Do: spread. For Route A, the lower half is \(12,13,14,14,15\), so \(Q_1=14\). The upper half is \(15,16,16,17,18\), so \(Q_3=16\). For Route B, the lower half is \(9,10,11,12,13\), giving \(Q_1=11\); the upper half is \(14,16,18,22,31\), giving \(Q_3=18\). Therefore:

$$ \text{IQR}_A=16-14=2\text{ minutes} \qquad \text{IQR}_B=18-11=7\text{ minutes} $$

The middle half of Route A’s times spans 2 minutes, compared with 7 minutes for Route B. Route B’s middle half is more spread out by \(7-2=5\) minutes. The ranges provide another view of the full spans: \(18-12=6\) minutes for Route A and \(31-9=22\) minutes for Route B. Range is sensitive to the extremes, so here the IQR is useful for comparing the middle half.

Do: outliers. Route A’s IQR fences are \(14-1.5(2)=11\) and \(16+1.5(2)=19\) minutes; all its values fall within those fences. Route B’s fences are \(11-1.5(7)=0.5\) and \(18+1.5(7)=28.5\) minutes. The value 31 is above 28.5, so it is flagged as a possible outlier by this rule. The rule flags an observation; it does not establish why the trip took that long.

Conclude in context. Route A’s times are fairly balanced around a median of 15 minutes, while Route B’s times are right-skewed with a median of 13.5 minutes. Route B has a wider middle half (IQR 7 minutes versus 2 minutes) and a high time of 31 minutes flagged by the 1.5 IQR rule. Thus, although Route A has the higher median, Route B’s observed travel times vary more, especially at the high end.

Why the Other Features Can Change the Story

The first example shows that the group with the higher median need not have the greater spread. It also shows why shape and outliers matter: a long upper tail or a flagged value can make one group’s distribution look quite different even when the centers are close. The next examples illustrate two other useful comparison habits: look beyond a small difference in center, and do not assume matching centers mean matching distributions.

Worked Example: Similar Medians, Different Distributions

Worked Example: Comparing Café Wait Times

A fictional café manager compares wait times at two service counters over 20 periods. The medians, IQRs, and graph descriptions are summarized below.

CounterShape shown by the histogramMedian waitIQRUnusual features
NorthApproximately symmetric8 minutes2 minutesNo apparent outliers
SouthRight-skewed7.5 minutes4.5 minutesOne isolated high wait of 19 minutes

Compare the wait-time distributions, using only what the graph descriptions and summaries support.

Shape: North’s wait times are approximately symmetric, while South’s distribution is right-skewed. The South histogram has a longer tail toward larger waiting times, including the isolated high wait. The shapes are not alike.

Center: North’s median wait is 8 minutes and South’s is 7.5 minutes. The medians differ by \(8-7.5=0.5\) minute, so North’s typical wait by this measure is only slightly longer. Because South is skewed and has an isolated high value, the median is a suitable center to compare, as in Choosing Mean or Median to Describe Center.

Spread: South’s IQR is 4.5 minutes, compared with 2 minutes for North. The middle half of South’s waits is therefore more spread out by \(4.5-2=2.5\) minutes. This is a comparison of the IQRs, not of the full ranges.

Outliers and conclusion: The graph description identifies an isolated high wait of 19 minutes for South and no apparent outliers for North. In context, the two counters have similar median waits, but South’s wait times are more variable in the middle half, are right-skewed, and include an isolated high observation. Reporting only the medians would miss those differences.

Worked Example: The Same Median Does Not Mean the Same Shape

Worked Example: Comparing Weekly Practice Times

Two fictional groups of students reported minutes of weekly practice for the same activity. The ordered observations are:

Group A (minutes)Group B (minutes)
2, 3, 4, 5, 6, 7, 8, 92, 2, 2, 5, 6, 9, 9, 9

Compare the distributions. Use the median and IQR for center and spread, and describe the shape and any outliers.

Center: Both groups have eight observations, so the median is the average of the fourth and fifth values. For each group:

$$ \text{Median}_A=\frac{5+6}{2}=5.5\text{ minutes} \qquad \text{Median}_B=\frac{5+6}{2}=5.5\text{ minutes} $$

The median practice time is the same in both groups. This does not make the full distributions alike.

Spread: For Group A, the lower half is \(2,3,4,5\), so \(Q_1=(3+4)/2=3.5\). The upper half is \(6,7,8,9\), so \(Q_3=(7+8)/2=7.5\). Its IQR is \(7.5-3.5=4\) minutes. For Group B, the lower half is \(2,2,2,5\), so \(Q_1=2\); the upper half is \(6,9,9,9\), so \(Q_3=9\). Its IQR is \(9-2=7\) minutes. Group B’s middle half spans 3 minutes more than Group A’s.

Shape and outliers: Group A has one observation at each value from 2 through 9, so its values are distributed relatively evenly across that span. Group B has concentrations at 2 and 9 with only two observations between them, suggesting a bimodal pattern. There is no clear isolated observation in either list. By the 1.5 IQR rule, Group A’s fences are \(3.5-1.5(4)=-2.5\) and \(7.5+1.5(4)=13.5\), and Group B’s are \(2-1.5(7)=-8.5\) and \(9+1.5(7)=19.5\). All observed values lie within their group’s fences.

Conclusion: Both groups have a median weekly practice time of 5.5 minutes, but Group B has a wider IQR and a more clearly bimodal pattern. Neither group has an observation flagged by the 1.5 IQR rule. In context, the shared median does not capture the difference in how practice times are distributed.

Common Mistakes and AP Exam Tips

  • Comparing only one statistic. Saying “Group A has a higher median” addresses center only. A complete comparison also describes shape, spread, and outliers for both groups.
  • Describing one group but not the other. “Route B is right-skewed” does not tell the reader what Route A looks like. Describe both groups, even when one has no apparent outliers or has a shape similar to the other.
  • Using vague words without evidence. “Group B is more spread out” is stronger when supported by a named measure: “Group B’s IQR is 7 minutes, compared with 2 minutes for Group A.” Include units and identify the variable.
  • Confusing spread measures. IQR describes the width of the middle half; range describes the distance from minimum to maximum; standard deviation describes a typical distance from the mean. Name the measure you are comparing and avoid treating them as interchangeable.
  • Calling a flagged value an error. The 1.5 IQR rule identifies a possible outlier. It does not prove the value is wrong or explain why it occurred. Describe what the data show; investigate the source only if information is available.
  • Claiming more than a graph supports. A boxplot can help compare medians, quartiles, and possible outliers, but it does not show every individual value or reliably reveal multiple peaks. Use the display that supports the claim.
  • Forgetting the context and units. “The IQR is larger” is less informative than “the middle half of South’s café waits spans 4.5 minutes, compared with 2 minutes at North.”

For full-credit communication, name the groups and variable, report relevant summaries with units, and compare the distributions feature by feature. Use cautious descriptions such as “approximately symmetric,” “suggests right skew,” or “no apparent outlier” when appropriate. Keep conclusions about the observed groups separate from claims about a larger population.

Key takeaway: A fair comparison does not stop at center. Describe shape, center, spread, and outliers for each group, then explain the similarities and differences using evidence and context.

Check Your Understanding

For each question, compare the groups in context and address all four distribution features when the information is available.

  1. Two groups’ quiz-score histograms are both approximately symmetric. Group A has median 78 points and IQR 8 points; Group B has median 82 points and IQR 15 points. State one comparison about center and one about spread.
  2. A boxplot shows a median of 12 minutes for Group X and 10 minutes for Group Y. What additional information would you need to compare their shapes, spreads, and outliers?
  3. Explain why two groups with the same median can still have different distributions. Use shape or spread in your explanation.
  4. A group’s IQR is 6 kilograms and another group’s range is 10 kilograms. Explain why those two numbers alone do not establish which group has greater spread.
  5. A value is above the upper fence in one group. What can you conclude from the 1.5 IQR rule, and what can you not conclude without more information?