Tutorials › AP Statistics › Choosing Class Width and Number of Bins

Graphs for quantitative data · Tutorial 67 of 1000

Choosing Class Width and Number of Bins

Learn how bin width and the number of bins affect a histogram’s appearance, and how to choose a useful scale for the data.

Beginner 9 min read

What You'll Learn

  • Explain how narrower and wider bins affect the detail shown in a histogram
  • Relate bin width to the number of bins needed to cover the data
  • Compare histograms made from the same data using different bin widths
  • Use consistent boundaries and equal-width bins to make counts comparable
  • Choose a bin width that shows useful patterns without obscuring the distribution

One Data Set, Different Histograms

In Constructing a Frequency Histogram, you learned to group quantitative values into consecutive bins and use the count in each bin as its bar height. But even when the data stay exactly the same, a histogram can look quite different if you change the bin width. A narrow width shows more detail; a wide width combines more observations in each bar.

There is usually no single required bin width or number of bins for a data set. The goal is to choose intervals that cover the observed values and make the distribution reasonably easy to see. Too few bins can hide features; too many can make the display jagged and difficult to summarize. Looking at more than one reasonable choice can help you judge whether a visible pattern is meaningful or an effect of the binning.

Definition: The bin width is the numerical length of each interval in a histogram. The number of bins is the number of intervals used to cover the data. For equal-width bins, a wider bin generally means fewer bins across the same span of values.

For equal-width bins, a useful relationship is:

$$ \text{number of bins} \approx \frac{\text{total span covered}}{\text{bin width}} $$

This is a planning relationship, not a rule that chooses the best histogram automatically. The span covered is determined by the first and last bin boundaries, which may extend beyond the minimum and maximum observations to use convenient endpoints. Make sure the bins cover every value, use a consistent boundary convention, and—when making a standard frequency histogram—keep the bin widths equal so that bar heights can be compared directly as counts.

What Changing the Width Does

With narrow bins, nearby values are separated into more intervals. This can reveal local concentrations, gaps, or changes in frequency, but small differences in the counts may make the histogram look uneven. With wide bins, neighboring intervals are combined. The bars tend to represent larger groups of observations, which can make the overall pattern easier to see but can conceal smaller features.

Changing the width does not change the observations or the total number of observations. It changes how those observations are grouped. If two adjacent bins are combined, the new frequency is the sum of their frequencies. This gives a useful check when comparing histograms: wider-bin counts should match the totals from the narrower bins they combine.

Choosing a width: Start with a width that gives a readable number of bins across the data range. Check that the intervals cover the minimum and maximum, and that the boundaries are easy to read. If the display seems too coarse or too irregular, try a different reasonable width and compare. Do not choose a width just because it makes one preferred pattern appear.

The examples below use one invented data set throughout. Each value is a completion time, in minutes, for one participant in a fictional puzzle activity. Comparing several widths makes it clear which features stay visible and which depend on the bin choice.

Worked Example: A Narrower Bin Width

Worked Example: A Narrower Bin Width

The 24 invented completion times, in minutes, are: 1, 2, 2, 3, 3, 3, 4, 4, 5, 5, 6, 6, 7, 7, 8, 9, 10, 10, 11, 12, 13, 14, 15, and 16. Construct a frequency table using bins of width 2 minutes, beginning at 0.

Set the intervals. Use \([0,2)\), \([2,4)\), \([4,6)\), \([6,8)\), \([8,10)\), \([10,12)\), \([12,14)\), and a final bin \([14,16]\) that includes 16. Each interval has width 2. As in the earlier tutorial on constructing a frequency histogram, a value at a shared boundary belongs to the bin that begins at that boundary.

Count and verify. The first bin contains 1, so its frequency is 1. The next contains two 2s and three 3s, for 5. The bins from 4 to less than 6 and 6 to less than 8 each contain two values of one integer and two of the next, giving 4 apiece. The remaining counts are 2, 3, 2, and 3. Adding the frequencies gives \(1+5+4+4+2+3+2+3=24\), matching the number of times listed.

Completion time (minutes)Frequency
0 to less than 21
2 to less than 45
4 to less than 64
6 to less than 84
8 to less than 102
10 to less than 123
12 to less than 142
14 through 16, including 163
Total24

Describe the effect of this width. The data are divided among eight bins, so this choice shows fairly fine detail. The most frequent bin is 2 to less than 4 minutes, with 5 observations. The later bins have smaller counts, though their frequencies vary from bin to bin. This view allows a reader to see those local changes, but it may also make small count differences look like notable features.

Worked Example: Combine Neighboring Bins

Worked Example: Combine Neighboring Bins

Use the same 24 completion times, but now make a frequency histogram with bins of width 4 minutes, beginning at 0. Compare its bin counts with the width-2 counts from the previous example.

Write the wider intervals. Use \([0,4)\), \([4,8)\), \([8,12)\), and a final bin \([12,16]\) that includes 16. Each bin combines two consecutive width-2 intervals from the previous table.

Add counts and check. For \([0,4)\), the observations 1, 2, 2, 3, 3, and 3 give a count of 6. For \([4,8)\), the two width-2 counts add to \(4+4=8\). For \([8,12)\), they add to \(2+3=5\). For \([12,16]\), they add to \(2+3=5\). The total is \(6+8+5+5=24\).

Completion time (minutes)FrequencyWidth-2 counts combined
0 to less than 461 + 5
4 to less than 884 + 4
8 to less than 1252 + 3
12 through 16, including 1652 + 3
Total24

Compare the appearance. The width-4 histogram has only four bars, so it gives a more compact summary. Its largest count is in the 4-to-less-than-8-minute interval. The table also shows why some detail has disappeared: the width-2 bars with counts 4 and 4 are now represented by one bar with count 8. The total is unchanged, but a reader can no longer distinguish the individual counts in those two narrower intervals.

Worked Example: A Wider Bin Width

Worked Example: A Wider Bin Width

Use the same data again, this time with bins of width 8 minutes. Begin at 0 and cover the values through 16.

Set two broad intervals. Use \([0,8)\) and a final bin \([8,16]\) that includes 16. Each bin is 8 minutes wide. To find their frequencies, add the counts from the four relevant width-2 bins.

Calculate the frequencies. The first interval combines the width-2 counts \(1+5+4+4=14\). The second combines \(2+3+2+3=10\). The total is \(14+10=24\), again matching the number of observations.

Completion time (minutes)FrequencyWidth-2 counts combined
0 to less than 8141 + 5 + 4 + 4
8 through 16, including 16102 + 3 + 2 + 3
Total24

Interpret the coarser display. With just two bars, the histogram gives a broad comparison: 14 completion times are below 8 minutes and 10 are from 8 through 16 minutes. This is easy to read, but it hides nearly all the variation within those two ranges. A two-bin histogram cannot show that the 2-to-less-than-4-minute interval had a higher count than its neighbors, or that the 8-to-less-than-12 and 12-through-16 intervals each had a count of 5.

Together, the three examples show the tradeoff. A narrow width can make local differences visible, a moderate width can summarize the distribution compactly, and a very wide width can hide important structure. All three histograms correctly represent the same 24 observations; they simply answer questions at different levels of detail.

A Practical Way to Choose

When choosing bins, think about the purpose of the display and the scale of the data. If the question concerns small changes in the values, a narrower width may be useful. If the main goal is a broad overview, a wider width may be clearer. A useful choice should produce enough bars to show the distribution’s shape without making the display overly crowded.

1
Find the data range.
Identify the minimum and maximum values and their units. Decide on convenient endpoints that will cover both.
2
Try a reasonable width.
Estimate how many equal-width bins will span the covered range. Choose boundaries that are simple to read and apply consistently.
3
Check the resulting display.
Look for a useful amount of detail. If there are very few bars, try a narrower width; if there are many bars with sparse counts, try a wider width.
4
Verify and describe.
Confirm that every observation is assigned once and that the counts add to the sample size. Describe the distribution shown by the selected bins, not a pattern that the graph cannot support.

There is no universal cutoff for what counts as “too many” or “too few” bins. The data context matters, and different reasonable choices can be useful for different purposes. If you compare two groups with histograms, use the same bin boundaries and width for both groups; otherwise, differences in the displays may come from the binning rather than the data.

Common Mistakes and AP Exam Tips

  • Assuming one width is automatically correct. A histogram is not wrong just because another reasonable width is possible. Explain what the chosen width helps show and what it may conceal.
  • Changing the boundaries as well as the width without noticing. A comparison is easiest to interpret when the bins begin at a common, sensible endpoint. State the intervals so readers can see exactly what is being counted.
  • Forgetting that wider-bin counts are sums. When adjacent bins are combined, add their frequencies. A wider-bin frequency should not be copied from just one of the smaller bins.
  • Ignoring the total count. No matter which width is used, the bin frequencies must add to the number of observations. If they do not, revisit the boundary assignments or tally.
  • Using different widths to compare groups. Use matching intervals when comparing histograms. Different bin widths can make bar heights and apparent patterns difficult to compare fairly.
  • Claiming that a feature is certain from one bin choice. A local peak or gap can depend on where the boundaries fall. When a claim matters, check whether it remains visible under another reasonable bin width.

For full-credit communication, name the bin width and intervals, explain how the frequencies were obtained, and verify the total. When comparing appearances, state what detail is visible with one width and what is combined or obscured with another. Describe counts in the units and context of the variable.

Key takeaway: Bin width controls the detail in a histogram. Narrow bins show more local variation; wider bins combine counts and can hide smaller features. Choose equal-width intervals that cover the data and make the distribution clear, and check that the frequency total stays the same.

Check Your Understanding

Use the ideas from the examples to reason about bin width and histogram appearance.

  1. If a histogram covering 0 through 20 uses bins of width 5, how many bins are needed?
  2. Two neighboring bins have frequencies 3 and 7. What is the frequency of the bin formed by combining them?
  3. What is one advantage and one disadvantage of using a wider bin width?
  4. Why should two groups being compared with histograms use the same bin boundaries?
  5. A histogram has bin frequencies that sum to 19, but the data set contains 20 observations. What should be checked?