Tutorials › AP Statistics › Describing Shape Center and Spread of a Distribution

Random variables and distributions · Tutorial 308 of 1000

Describing Shape Center and Spread of a Distribution

Use a probability histogram to describe a distribution’s overall pattern, locate its approximate center, and compare how concentrated or spread out its probabilities are.

Intermediate 8 min read

What You'll Learn

  • Describe a probability histogram as roughly symmetric, skewed, or multimodal.
  • Distinguish a distribution’s peak from its approximate center.
  • Estimate a probability-weighted center from values and probabilities.
  • Describe spread by considering the range and where probability is concentrated.
  • Compare distributions by discussing shape, center, and spread separately.

Look Beyond the Tallest Bar

In Drawing a Probability Histogram, you learned that each bar is centered at a possible value of the random variable and that its height represents the probability of that value. Now use the whole pattern of bars to describe a distribution. Three useful features are its shape, center, and spread.

The tallest bar identifies the most likely single value, also called the mode. But it does not, by itself, identify the center. To judge the center, consider the locations of all the bars and how much probability lies at each location. A bar with probability 0.30 contributes more to the overall balance than a bar with probability 0.04.

Definition: The shape describes the overall pattern of probabilities in a distribution. The center describes a typical or balancing location for the distribution. The spread describes how far the possible values extend and how concentrated or dispersed their probabilities are.

For a discrete distribution, one way to describe its center is the probability-weighted center: values with larger probabilities have more influence on the balance. In this tutorial, we will estimate that center by multiplying each possible value by its probability and adding the products. This is more informative than simply choosing the tallest bar.

$$ \text{Probability-weighted center} =\sum xP(X=x) $$

This calculation gives a useful numerical center, but the histogram remains important. A single center does not show whether the distribution is symmetric, skewed, or has more than one peak. Likewise, spread is not captured by the center alone.

Describe the Shape

Start by looking at the heights and arrangement of the bars. A distribution is unimodal when it has one clear peak, and bimodal when it has two distinct peaks. A distribution with bars of roughly similar heights across its values may be described as approximately uniform. These words describe the overall pattern, not just one bar.

A distribution is approximately symmetric when the bars on either side of its center have roughly similar heights at corresponding distances. If the distribution has a longer or heavier tail toward larger values, it is skewed right. If the longer or heavier tail extends toward smaller values, it is skewed left. In a skewed distribution, the peak and center can be noticeably different.

The terms “right” and “left” refer to the direction of the tail along the numerical axis. For example, a distribution can peak at 2 but have more probability at values above 2 than at values below 2. That pattern is not made symmetric just because the tallest bar is at 2.

Key takeaway: Describe the full arrangement of bars. Identify the number of peaks and whether the bars are roughly balanced or have a tail extending toward smaller or larger values.

Estimate the Center and Describe the Spread

A quick visual estimate of the center asks where the distribution would balance if each possible value were a location and its probability acted like weight. For a rough description, locate the region where the probability is concentrated, then check how much probability lies on either side. If probabilities are given in a table, the probability-weighted calculation makes that balance more precise.

For spread, first note the smallest and largest possible values. Their difference is the range of possible values. Then look between those endpoints: a distribution with most of its probability near the center is more concentrated than one that places substantial probability far from the center. Two distributions can have the same range but very different concentration, so the endpoints alone do not tell the whole story.

Conditions: To describe a distribution from its probability histogram, use the numerical horizontal scale, consider every bar rather than only the tallest one, and distinguish the overall range from the concentration of probability within that range.

Worked Example: Describe the Number of Children

Worked Example: Describe the Number of Children

For an invented chance model, let \(X\) be the number of children in a randomly selected household from a defined community. Suppose its probability distribution is:

Number of children, \(x\)\(P(X=x)\)
00.10
10.25
20.35
30.20
40.10

State. Describe the shape, approximate center, and spread of the probability histogram for \(X\).

Plan. Use the relative bar heights to describe shape, and compare probabilities at values on either side of the peak. Estimate the center using the probability-weighted calculation. Describe spread by noting the possible-value range and where most of the probability lies.

Do. The tallest bar is at 2 children, with probability 0.35. The bars rise to that value and generally fall afterward, so the distribution has one clear peak. It is roughly symmetric, though not perfectly: the probabilities at 0 and 4 match at 0.10, while the probability at 1 is slightly larger than the probability at 3 (0.25 versus 0.20).

The probability-weighted center is \(0(0.10)+1(0.25)+2(0.35)+3(0.20)+4(0.10)=1.95\). Thus, the center is approximately 2 children. The range of possible values is from 0 to 4, a range of 4 children. Most of the probability is at 1, 2, or 3 children: \(0.25+0.35+0.20=0.80\).

$$ \text{Center} =0(0.10)+1(0.25)+2(0.35)+3(0.20)+4(0.10) =1.95 $$

Conclude. The distribution is unimodal and roughly symmetric, with its peak at 2 children and its approximate center at 1.95, or about 2 children. Its possible values range from 0 to 4, with 80% of the probability between 1 and 3 children, inclusive.

A Peak Is Not Necessarily the Center

A distribution’s most likely value and its center answer different questions. The peak identifies the single value with the greatest probability. The center reflects the balance of probability across all possible values. The distinction matters especially when the probabilities are not arranged symmetrically around the peak.

Worked Example: Compare Two Message-Count Distributions

Consider two invented models for \(X\), the number of messages a device receives during a specified period. Model A and Model B have these distributions:

Messages, \(x\)Model AModel B
00.100.04
10.200.20
20.400.30
30.200.20
40.100.20
5—0.06

State. Compare the peaks, approximate centers, and patterns of the two distributions.

Plan. First find the tallest bar in each model. Then calculate the probability-weighted center for each, so the comparison uses all possible values rather than just the peaks. Finally, compare the probabilities on the lower and higher sides.

Do. Both models have their tallest bar at 2 messages. For Model A, the center is \(0(0.10)+1(0.20)+2(0.40)+3(0.20)+4(0.10)=2.00\). The bars are symmetric around 2: the probabilities at 0 and 4 match, as do those at 1 and 3.

For Model B, the center is \(0(0.04)+1(0.20)+2(0.30)+3(0.20)+4(0.20)+5(0.06)=2.50\). Although its peak is also at 2, Model B has more probability on the high side: for example, \(P(X=4)=0.20\), compared with \(P(X=0)=0.04\). Its probabilities are not balanced around the peak and extend to 5 messages.

$$ \begin{aligned} \text{Model A center}&=0.00+0.20+0.80+0.60+0.40=2.00\\ \text{Model B center}&=0.00+0.20+0.60+0.60+0.80+0.30=2.50 \end{aligned} $$

Conclude. Both distributions peak at 2 messages, but their centers differ: Model A is centered at 2, while Model B is centered at 2.5 and has more probability on the high side. The peak alone would miss that shift.

Same Center, Different Spread

The center also does not tell you how concentrated a distribution is. Imagine two histograms balanced around 2. One might place most of its probability on values close to 2; the other might place much more probability at values farther away. Both can have the same center while having different spreads.

Worked Example: Compare Concentration Around the Same Center

Suppose two invented models describe \(Y\), the number of pieces needing adjustment in a small batch. Their possible values and probabilities are:

Pieces needing adjustment, \(y\)Model CModel D
00.020.10
10.080.20
20.800.40
30.080.20
40.020.10

State. Compare the center and spread of the two models.

Plan. Calculate each probability-weighted center. Then compare how much probability each model puts near 2 and at the more distant values 0 and 4.

Do. Model C has center \(0(0.02)+1(0.08)+2(0.80)+3(0.08)+4(0.02)=2.00\). Model D has center \(0(0.10)+1(0.20)+2(0.40)+3(0.20)+4(0.10)=2.00\). Both are symmetric around 2 and have the same possible-value range, from 0 to 4. But Model C places 0.80 probability at 2, whereas Model D places only 0.40 there. Model C puts just \(0.02+0.02=0.04\) at the endpoints, while Model D puts \(0.10+0.10=0.20\) at the endpoints.

$$ \begin{aligned} \text{Model C center}&=0.08+1.60+0.24+0.08=2.00\\ \text{Model D center}&=0.20+0.80+0.60+0.40=2.00 \end{aligned} $$

Conclude. Both models are centered at 2 and have the same range, but Model C is more concentrated near its center. Model D is more spread out because it assigns more probability to values away from 2.

Common Mistakes and AP Exam Tips

  • Calling the mode the center. The tallest bar gives the most likely single value, not necessarily the center. A full-credit answer distinguishes the peak from the approximate balance point.
  • Describing shape from one bar. A distribution with a peak at 2 is not automatically symmetric or centered at 2. Compare probabilities on both sides and consider the full numerical range.
  • Reversing the direction of skew. Right-skewed means the tail extends toward larger values; left-skewed means it extends toward smaller values. Name the direction of the tail, not the side with the peak.
  • Using range as the only measure of spread. The range identifies the endpoints, but distributions with the same endpoints can have very different concentrations. Describe where the probability is located as well.
  • Ignoring probability weights when estimating center. A possible value with probability 0.30 should influence the balance more than one with probability 0.04. Use all values and their probabilities for the weighted calculation.

For a clear AP response, name the shape and support the description with the pattern of bars. Give the center in context and do not infer it from the peak alone. For spread, mention the possible-value range and whether probabilities cluster near the center or extend farther away.

Key Takeaway

A probability histogram tells more than which value is most likely. Read its full pattern to describe shape, use all the values and probabilities to judge its approximate center, and examine both its range and concentration to describe spread.

Key takeaway: Keep shape, center, and spread distinct. A peak is the most likely value, not necessarily the center; and the range alone does not describe how concentrated the probability is.

Check Your Understanding

Use the distribution below for \(Z\), the number of reusable containers returned to a collection station during a short interval.

Containers returned, \(z\)\(P(Z=z)\)
00.10
10.20
20.40
30.20
40.10
  1. Describe the shape of the distribution and identify its peak.
  2. Calculate the probability-weighted center of \(Z\).
  3. What is the range of possible values?
  4. Describe where most of the probability is concentrated.
  5. Explain why knowing the peak alone is not enough to describe the center and spread of a distribution.