Tutorials › AP Statistics › Overlap and Separation Between Two Distributions

Comparing distributions · Tutorial 133 of 1000

Overlap and Separation Between Two Distributions

Learn to describe how much two distributions overlap, where they separate, and what that comparison does—and does not—suggest about group differences.

Beginner 9 min read

What You'll Learn

  • Compare overall ranges and identify the numerical interval shared by two groups.
  • Use the overlap between boxplot boxes to describe how their middle halves compare.
  • Count observations in a shared interval when raw data or dotplots are available.
  • Distinguish overlap in possible values from similarity of the distributions.
  • Explain what overlap suggests about separating individual observations by group.
  • Avoid treating visual overlap as a test of statistical significance or a cause-and-effect conclusion.

Overlap Adds Another View of Group Differences

In Comparing Distributions, you learned to compare shape, center, spread, and unusual features. Overlap adds a useful question: across what values do the groups share observations, and where are their values more distinct? Two groups can have different centers while still containing many observations in a shared range. Conversely, a modest difference in centers can look more pronounced when each group has little spread.

Overlap is not a single summary statistic with one universally accepted calculation. It is a way to describe how two distributions line up on a common scale. The display matters: dotplots and histograms show where observations or intervals are concentrated, while side-by-side boxplots show the ranges and middle halves. As in Comparing Two Histograms on the Same Scale and Side-by-Side Boxplots Compared, compare displays only when their scales are aligned.

Definition: Two distributions overlap when both have observations in some of the same numerical region. Separation describes how distinct their values or concentrations appear. The amount of overlap depends on which part of the distributions is being considered and what the display shows.

A shared range is a starting point, not a complete description. The minimum-to-maximum ranges may overlap even if most observations cluster apart. Boxplot boxes may overlap only a little, even when whiskers extend through a much wider shared range. A histogram or dotplot can reveal whether observations are concentrated in the shared region or only a few values reach it.

A practical comparison uses three layers: first check the full ranges, then compare the central regions, and finally inspect the shape and concentration of the observations in any shared region. State which layer supports your conclusion. “The ranges overlap” is a specific claim; “the groups are almost the same” is much broader and may not be supported.

What Overlap Can Suggest

When two groups overlap substantially, a value from one group may also be quite plausible in the other. Knowing an individual observation’s value may therefore not clearly identify which group it came from. When distributions are more separated, values may be more useful for distinguishing the groups—but overlap and separation alone do not guarantee that every observation can be classified correctly.

The comparison is about distributions, not a rule for every individual. Even when one group generally has larger values, some observations in that group may be below observations in the other group. Say that a group’s values tend to be higher or that its distribution is shifted higher when the display supports that description. Avoid saying that every member of one group has a higher value.

Overlap also does not determine whether an observed group difference is statistically significant. That question requires an appropriate inference procedure and information about how the data were collected. Nor does a difference between groups by itself show that group membership caused the difference. In this tutorial, use overlap as a descriptive feature alongside shape, center, spread, and unusual values.

Key idea: Describe overlap at the level shown by the display: full ranges, middle halves, or concentrations of observations. Then explain cautiously what that pattern suggests about how distinct the groups look.

Comparing Shared Ranges

For a quick first check, identify each group’s minimum and maximum, then find the numerical interval that both ranges cover. If one group’s range is from 14 to 27 centimeters and the other’s is from 18 to 32 centimeters, both ranges include values from 18 to 27 centimeters. The ranges overlap on that interval.

The width of a shared interval can help describe its size in the variable’s units. In this example, the interval’s width is \(27-18=9\) centimeters. But that width is not a percentage of the observations that overlap, and it does not show how densely either group’s observations fall in the interval. For that, inspect the individual observations or a suitable graph.

Be careful with endpoints and discrete data. If measurements are recorded as whole numbers, an interval from 18 to 27 includes those recorded values, but the arithmetic width is 9 centimeters. That is different from counting the 10 possible whole-number values 18, 19, and so on through 27. Say whether you are describing interval width or counting recorded values.

Worked Example: Comparing the Ranges of Two Plant Groups

A fictional greenhouse records the heights, in centimeters, of 12 seedlings from each of two growing groups. The observations are listed in increasing order:

Group A: 14, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 27
Group B: 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 30, 32

Question: Do the groups’ full ranges overlap, and what does that observation suggest?

Find the ranges. Group A extends from 14 to 27 centimeters. Group B extends from 18 to 32 centimeters. The interval shared by these full ranges begins at the larger minimum, 18, and ends at the smaller maximum, 27.

$$ \text{Shared range}=[18,27]\text{ cm}, \qquad \text{width}=27-18=9\text{ cm}. $$

The arithmetic checks by subtracting the shared interval’s endpoints: \(27-18=9\). Both groups have observations throughout much of this shared region; for instance, Group A has values from 18 through 27, and Group B also has values from 18 through 27, though not every whole-number value appears in each group.

Conclude. The full ranges overlap from 18 to 27 centimeters, so the groups are not completely separated by height. The ranges alone do not tell us how much of each group is concentrated in that interval or whether their distributions otherwise look similar. A dotplot or histogram would provide more detail about those patterns.

Overlap Between the Middle Halves

A side-by-side boxplot offers another view. Each box extends from \(Q_1\) to \(Q_3\), so it represents the middle half of that group’s observations. If the boxes overlap, the middle halves share a numerical interval. If the boxes barely touch or do not overlap, the central portions appear more separated, even if the whiskers share values.

The box overlap is about the interval shown by the boxes. It is not a count of how many observations from each group are in that interval: a boxplot does not show the individual data values. And a box that is twice as wide as another does not mean it contains more observations. Each box represents the middle half of its own group.

As in Comparing Spread Using IQR and Standard Deviation, compare matching measures when discussing spread. The IQR describes the width of a box; the overlap between boxes describes how their middle-half intervals line up. These are related but different observations. You can report both without treating one as a substitute for the other.

Worked Example: Reading Overlap Between Two Boxplot Boxes

A fictional park compares the time, in minutes, that visitors spend at two nature exhibits. Their side-by-side boxplots have these summaries:

ExhibitMinimum\(Q_1\)Median\(Q_3\)Maximum
Creek3542485463
Meadow4350566270

Question: How do the full ranges compare with the overlap between the middle halves?

Compare the full ranges. The Creek range is 35 to 63 minutes, and the Meadow range is 43 to 70 minutes. Their shared interval is 43 to 63 minutes, with width \(63-43=20\) minutes.

Compare the boxes. The Creek box extends from \(Q_1=42\) to \(Q_3=54\). The Meadow box extends from \(Q_1=50\) to \(Q_3=62\). The shared interval between these boxes is 50 to 54 minutes, with width \(54-50=4\) minutes. Each box has width \(54-42=12\) minutes and \(62-50=12\) minutes, respectively. Thus, the overlapping span is \(4/12=1/3\) of each box’s width.

That one-third calculation compares interval widths; it does not mean that exactly one-third of the visitors at either exhibit spent time in the shared interval. A boxplot does not show the individual values needed to count that share.

Conclude. The full ranges overlap considerably, but the middle-half boxes overlap only from 50 to 54 minutes. The medians are 48 minutes at Creek and 56 minutes at Meadow, so the Meadow distribution has a higher median. In context, the central parts of the two time distributions show some overlap, but the boxplots also suggest a shift toward longer visits at Meadow. The displays do not establish that every Meadow visitor stays longer.

Using Individual Values to Describe a Shared Region

When raw data or a dotplot is available, you can count how many observations from each group fall in a shared numerical interval. This makes the description more specific than range overlap alone. State the interval, count observations in it for each group, and, if useful, report each count as a fraction or percentage of that group.

This calculation answers a limited question: what proportion of each group’s observed values lies in the specified interval? It does not by itself measure every aspect of overlap. Two distributions might put the same proportion in an interval but arrange those observations differently within it. Describe any visible clustering or gaps as well.

Worked Example: Counting Values in a Shared Score Interval

A fictional recreation program records the number of laps completed by participants during a practice session. Two groups each have 10 participants:

Group Pine: 5, 6, 7, 8, 9, 10, 11, 12, 13, 14
Group Birch: 8, 9, 10, 11, 12, 13, 15, 16, 17, 18

Question: What share of each group’s observations falls in the shared interval of the two full ranges?

Identify the shared interval. Pine’s range is 5 to 14 laps, and Birch’s range is 8 to 18 laps. The ranges share the interval from 8 to 14 laps.

Count Pine observations in the interval. Pine has 8, 9, 10, 11, 12, 13, and 14 laps in the interval: 7 observations out of 10. The percentage is

$$ \frac{7}{10}\times 100\%=70\%. $$

This checks because \(7\div10=0.7\), or 70%. Birch has 8, 9, 10, 11, 12, and 13 laps in the interval: 6 observations out of 10. Its percentage is

$$ \frac{6}{10}\times 100\%=60\%. $$

The second calculation checks because \(6\div10=0.6\), or 60%.

Conclude. The full ranges overlap from 8 to 14 laps. Seven of Pine’s 10 observations and six of Birch’s 10 observations lie in that shared interval. The groups therefore have many values in a common region, although Pine also has lower values and Birch has higher values outside it. This is a description of these observations, not proof that the groups will have the same pattern in a broader population.

Common Mistakes and AP Exam Tips

  • Treating range overlap as complete similarity. Overlapping minimum-to-maximum intervals do not show where most values lie. Look at the boxes or the shape and concentration in a dotplot or histogram.
  • Confusing the width of an interval with the number of observations in it. Subtracting endpoints gives a width in the variable’s units. Counting observations requires raw data or a display that shows them.
  • Calling box overlap a percentage of observations. Boxplots show quartiles, not individual values within the box. Describe the shared interval between the boxes, not an unsupported count or percentage of observations.
  • Assuming overlapping groups have identical centers. The medians or means may still differ. Compare the centers directly, then describe the overlap as a separate feature.
  • Claiming no individuals can be confused when groups look separated. Even an apparent separation in a sample does not guarantee separation for every person or in a larger population. Keep the conclusion tied to the observations shown.
  • Turning a descriptive pattern into an inference or causal claim. Overlap does not provide a p-value, establish statistical significance, or show that one group characteristic caused higher or lower values.

A full-credit comparison names both groups and the variable, identifies the specific region that overlaps or appears separated, and supports the statement with values from the display. For example: “The middle-half intervals overlap from 50 to 54 minutes, although the Meadow median is 8 minutes higher than the Creek median.” That is more precise than saying only that the groups are “different” or “about the same.”

Key takeaway: Judge overlap at a stated level—full ranges, boxplot boxes, or observed concentrations. Explain what the shared and separated regions suggest about group differences, but do not treat descriptive overlap as a test of significance or a claim about cause.

Check Your Understanding

Use the shared scale and the specific display feature named in each question.

  1. Group A ranges from 12 to 28 units, and Group B ranges from 20 to 35 units. What interval do their full ranges share, and what is its width?
  2. Two boxplot boxes extend from 30 to 44 and from 40 to 54. State the interval where the middle halves overlap. What does the box overlap not tell you?
  3. Why does overlap between full ranges not necessarily mean that most observations are in a shared region?
  4. One group’s median is higher than another’s, but their boxes overlap. Write one cautious sentence describing both features.
  5. Why can’t a visual judgment about overlap alone establish statistical significance or cause and effect?