Tutorials › AP Statistics › Comparing Two Probability Distributions

Random variables and distributions · Tutorial 317 of 1000

Comparing Two Probability Distributions

Compare defect-count distributions by looking at their typical values and their concentration or variation, not just at their tallest bars.

Intermediate 9 min read

What You'll Learn

  • Describe the center of each defect-count distribution in context.
  • Compare spread using both the range and how concentrated values are near the center.
  • Use relative frequencies to compare histograms when the sample sizes differ.
  • Explain why equal ranges do not necessarily mean equal spread.
  • Write a direct comparison that identifies which machine tends to produce more defects and which has more variable counts.

Comparing Distributions Side by Side

In Describing Shape Center and Spread of a Distribution, you learned to describe the shape, center, and spread of one distribution. Comparing two distributions uses those same ideas, but adds a direct question: how are the patterns alike, and how do they differ? Here we will compare histograms of the number of defects on items produced by two machines.

Suppose quality-control staff inspect a random sample of items from each machine and record the number of defects on each item. Let \(X_A\) be the number of defects on an item from Machine A, and let \(X_B\) be the number on an item from Machine B. A histogram places the defect count on the horizontal axis and shows how often each count occurred. When the samples are equally large, counts can be compared directly. If the sample sizes differ, relative frequencies—the proportions of inspected items at each count—make a fairer comparison.

Definition: To compare two distributions, describe each distribution’s shape, center, and spread, then state the important differences in context. The center describes a typical or balancing location; the spread describes how much the values vary or how concentrated they are around that center.

A comparison of center asks where the distributions are located on the horizontal axis. If a typical item from Machine B has more defects than a typical item from Machine A, Machine B’s distribution is centered farther to the right. The tallest bar can help locate a peak, but it is not automatically the center. As in the earlier tutorial on distribution shape, center and most likely value are not interchangeable.

A comparison of spread asks which distribution has more variation. The range—the largest observed value minus the smallest—is one useful clue, but it is not the whole comparison. Two distributions can have the same range while one places most observations near its center and the other places more observations farther away. Look at the full pattern of bars, including how much is concentrated near typical values and how much is in the tails.

A Method for Comparing Two Histograms

Before describing the distributions, check that the histograms use the same horizontal scale and the same bins. Otherwise, a difference in the display could make a comparison misleading. Also identify the sample size for each machine. When sample sizes differ, compare relative frequencies rather than raw bar heights; for example, 12 items out of 40 is 30%, while 12 items out of 60 is 20%.

1
Identify what each distribution represents.
Name the measured variable and connect each histogram to its machine and group of inspected items.
2
Compare centers.
Use the horizontal locations of the bulk of the data or the medians, if available, to say which machine’s defect counts are typically higher or lower.
3
Compare spreads.
Compare ranges, then look at how concentrated the bars are near the center and how much lies farther away.
4
State the contrast in context.
Say which machine has the higher typical defect count, which has more variation, or whether a particular feature is similar. Support each claim with details from the distributions.

For a discrete defect count, a median can be found by ordering the observations and locating the middle value or values. With an even number of observations, the median is the average of the two middle values. The median gives one way to make “typical” precise, while the histogram still shows features the median alone cannot describe. For example, it does not show whether values are clustered tightly or spread across the range.

Worked Example: One Machine Has a Higher Typical Count

In an invented quality-control exercise, staff inspect 40 items from each machine. The table gives the number of items observed at each defect count. Let \(X_A\) and \(X_B\) be the number of defects on one inspected item from Machine A and Machine B, respectively.

Defects, \(x\)Machine A countMachine A relative frequencyMachine B countMachine B relative frequency
080.2020.05
1160.4080.20
2120.30160.40
340.10100.25
400.0040.10

State. Compare the centers and spreads of the defect-count distributions for the inspected items.

Plan. Each sample contains 40 items, so the counts are directly comparable. Find the medians to describe a typical count, compare the observed ranges, and use the table to check where the observations are concentrated.

Do. For Machine A, the ordered observations in positions 1 through 8 have 0 defects, and positions 9 through 24 have 1 defect. The 20th and 21st observations are both 1, so the median is \((1+1)/2=1\) defect. For Machine B, positions 1 through 2 have 0 defects, positions 3 through 10 have 1, and positions 11 through 26 have 2. The 20th and 21st observations are both 2, so the median is \((2+2)/2=2\) defects.

Machine A’s observed range is \(3-0=3\) defects. Machine B’s observed range is \(4-0=4\) defects. The histograms would also show that Machine A’s counts are concentrated from 0 to 2 defects: \(8+16+12=36\) of 40 items, or \(36/40=0.90\). Machine B has 2+8+16=26 items from 0 to 2 defects, or \(26/40=0.65\); more of its observations extend to 3 or 4 defects.

Conclude. Among these inspected items, Machine B has a higher typical defect count: its median is 2 defects, compared with 1 defect for Machine A. Machine B also shows greater spread, with a larger range and more observations extending to higher counts. This describes the observed distributions; it does not by itself establish why the machines differ or guarantee the counts on future items.

Same Center, Different Spread

A comparison can show little or no difference in center even when the distributions differ substantially in spread. For a clear comparison, choose a neighborhood around the shared center and count or estimate what proportion of observations falls inside it. Then compare how much lies farther from the center. This complements the range: range focuses only on the two extreme observations, while concentration considers many observations.

The idea of a “neighborhood” should match the scale and context. For defect counts with a median of 2, for instance, values from 1 through 3 are within one defect of the median. A higher proportion in that interval suggests greater concentration near the center. It is still useful to mention other features, such as unusually low or high counts, rather than relying on one interval alone.

Worked Example: Equal Medians but Unequal Concentration

In a second invented exercise, 40 items from each machine are inspected. The frequency table represents the bars in two histograms on the same scale.

Defects, \(x\)Machine A countMachine B count
0412
1124
2168
384
4012

State. Compare the centers and spreads of the two distributions.

Plan. Find the medians, compare the ranges, and then compare the proportions within one defect of the median. This last comparison checks concentration near the center rather than only the endpoints.

Do. For Machine A, the 20th and 21st observations are both 2 defects, so its median is 2. For Machine B, positions 1 through 12 have 0 defects, 13 through 16 have 1, and 17 through 24 have 2. Its 20th and 21st observations are also both 2, so its median is 2. Both distributions have range \(3-0=3\) for A and \(4-0=4\) for B. To compare concentration within one defect of the median, count observations with 1, 2, or 3 defects. Machine A has \(12+16+8=36\), which is \(36/40=0.90\). Machine B has \(4+8+4=16\), which is \(16/40=0.40\).

Conclude. The two inspected groups have the same median defect count, 2, but Machine A’s counts are much more concentrated within one defect of that median. Machine B’s distribution is more spread out, including many items at 0 and 4 defects. The medians alone would miss this difference.

Equal Ranges Can Hide Differences

Equal ranges do not mean that two distributions have the same spread. The endpoints tell you only the smallest and largest observed values. They do not tell you how many observations are at those endpoints, how far apart the other values are, or whether the observations cluster near a typical value. Histograms make these features visible, so use the full distribution when making a spread comparison.

It is also possible for distributions to have the same range but different centers. When one histogram is generally shifted to the right, its typical count is higher, even if both histograms extend over the same number of defect-count values. A careful comparison treats center and spread as separate features rather than using one to stand in for the other.

Worked Example: Same Range, Different Center and Tail Pattern

For a third invented quality-control exercise, 40 items are inspected from each machine. Both machines have some items with 0 defects and some with 4 defects.

Defects, \(x\)Machine A countMachine B count
028
1816
2208
384
424

State. Compare the distributions’ centers and spreads, taking care not to rely on range alone.

Plan. Use the medians to compare centers. Calculate each range, then compare the counts within one defect of each median and note where the observations fall.

Do. For Machine A, positions 1 and 2 have 0 defects, positions 3 through 10 have 1, and positions 11 through 30 have 2. Its 20th and 21st observations are both 2, giving a median of 2 defects. For Machine B, positions 1 through 8 have 0 defects and positions 9 through 24 have 1, so its 20th and 21st observations are both 1. Its median is 1 defect. Both ranges are \(4-0=4\) defects. Within one defect of the median, Machine A has 8+20+8=36 items from 1 to 3 defects, or \(36/40=0.90\). Machine B has 8+16+8=32 items from 0 to 2 defects, or \(32/40=0.80\). Machine B also has \(8+4=12\) items at the endpoints 0 or 4, compared with \(2+2=4\) for Machine A.

Conclude. Machine A’s inspected items have a higher median defect count than Machine B’s: 2 rather than 1. Although the ranges are equal, Machine B has more observations at the extremes and a smaller proportion within one defect of its median, so its distribution is less concentrated near its center. The same range does not make the distributions equally spread out.

Common Mistakes and AP Exam Tips

  • Reporting one distribution instead of comparing both. A response that says “Machine A’s counts range from 0 to 3” does not state how that compares with Machine B. Make the contrast explicit: “Machine B’s range is one defect wider.”
  • Calling the tallest bar the center. The tallest bar identifies the most frequent value, not necessarily a typical or balancing location for the whole distribution. Describe the overall location or use medians when the data are available.
  • Using range as the only evidence about spread. Range can be affected by just two observations. Also describe how concentrated the values are and whether many observations lie far from the center.
  • Comparing counts when sample sizes differ. A larger sample can have taller bars simply because more items were inspected. Use relative frequencies when the group sizes are unequal.
  • Ignoring the direction and context. “The distributions differ” is not specific enough. Say which machine has the higher typical number of defects or the greater variation, and include the defect-count units.
  • Making a claim beyond the data. A comparison of inspected items describes those distributions. Do not claim that the machine itself caused a difference or that every future item will follow the same pattern without additional evidence.

A strong AP response names the variable, compares center and spread separately, and supports each comparison with an appropriate feature such as the median, range, or concentration near the center. It ends with a sentence in context—for example, “Machine B’s inspected items have a higher typical defect count, while Machine A’s counts are more concentrated near their median.”

Key Takeaway

Compare two distributions by asking where each is centered and how much its values vary. Use the full histogram: a range is helpful, but the concentration of observations near the center and the amount in the tails matter too.

Key takeaway: Describe center and spread as separate features, then state the comparison in context. Equal ranges do not guarantee equal spread, and the tallest bar is not automatically the center.

Check Your Understanding

Use the defect counts below to compare the distributions for two invented machines. Each machine has 20 inspected items.

Defects, \(x\)Machine A countMachine B count
026
184
286
324
  1. Find the median defect count for each machine. What do the medians suggest about their centers?
  2. Find each machine’s observed range. Does this alone show which distribution is more concentrated near its center?
  3. For each machine, find the proportion of items with 1 or 2 defects. What does this tell you about concentration in those values?
  4. Write a two-sentence comparison of center and spread in context.
  5. Why would comparing raw bar heights be misleading if one machine had 20 inspected items and the other had 50?