Reading the Shape of a Binomial Distribution
In Standard Deviation of a Binomial Distribution, you learned that for \(X\sim B(n,p)\), the mean is \(np\) and the standard deviation is \(\sqrt{np(1-p)}\). Those describe the center and typical spread. A histogram adds another feature: shape, or the way the probabilities are distributed across the possible counts.
A binomial histogram has one bar for each possible number of successes, from 0 through \(n\). The height of the bar at \(k\) represents \(P(X=k)\). The pattern of bar heights can be balanced around the center, concentrated near one end with a tail toward the other, or somewhere between these patterns.
For a binomial model, the value of \(p\) gives a useful first clue about shape. When \(p=0.5\), success and failure are equally likely, so the binomial probabilities are symmetric. When \(p\) is below 0.5, the histogram tends to have more probability at lower counts and a tail toward higher counts. When \(p\) is above 0.5, the pattern tends to reverse.
These are useful shape descriptions, not substitutes for looking at the probabilities. The value of \(n\) also matters. For a fixed \(p\), increasing \(n\) changes the possible counts and the distribution’s center and spread. Its histogram may look less strongly skewed relative to its center, but it is not necessarily symmetric. In this tutorial, use the probability pattern to support a shape description rather than relying only on a rule of thumb.
How the Parameters Guide Your First Impression
The mean, \(np\), helps locate the center of a binomial distribution, but it does not have to be a possible count. For example, a mean of 2.5 lies between the possible counts 2 and 3. The mean also does not have to match the tallest bar. A histogram’s shape depends on all its probabilities, not just its center.
The standard deviation, \(\sqrt{np(1-p)}\), describes typical distance from the mean, as explained in the previous tutorial. It does not tell you whether the distribution is symmetric or skewed. Two distributions can have similar standard deviations but different shapes, just as two distributions can have the same general shape but different centers and spreads.
As in Recognizing a Binomial Setting and The BINS Checklist for Binomial Conditions, first identify what one trial is, what counts as success, and what \(X\) counts. The comparisons below assume that each model satisfies the binomial conditions. For a specific real-world process, those conditions still need justification.
Worked Example: A Symmetric Distribution
Worked Example: A Symmetric Distribution
A game has 4 independent rounds. In each round, a player has probability 0.5 of winning. Let \(X\) be the number of rounds the player wins. Describe the shape of the distribution.
State. The count is \(X\sim B(4,0.5)\). The trial is one round, success is winning that round, and \(X\) counts wins.
Plan. Each round has two outcomes, the number of rounds is fixed at 4, the rounds are independent, and the probability of a win is the same, 0.5, in every round. These are the binomial conditions. To judge shape, calculate the probabilities for each possible count and compare values equally far from the center.
Do. Use the binomial probability formula for \(k=0,1,2,3,4\):
| Wins \(k\) | \(P(X=k)\) |
|---|---|
| 0 | 0.0625 |
| 1 | 0.2500 |
| 2 | 0.3750 |
| 3 | 0.2500 |
| 4 | 0.0625 |
For example, \(P(X=1)=\binom{4}{1}(0.5)^1(0.5)^3=4(0.0625)=0.2500\). The probabilities sum to \(0.0625+0.2500+0.3750+0.2500+0.0625=1.0000\). Counts equally far from 2 have equal probabilities: \(P(X=0)=P(X=4)\) and \(P(X=1)=P(X=3)\).
Conclude. The histogram is symmetric around 2 wins. This matches the mean \(np=4(0.5)=2\). The bars rise toward 2 and then fall in a mirror-image pattern.
Worked Example: A Right-Skewed Distribution
Worked Example: A Right-Skewed Distribution
A wildlife camera records whether a certain animal appears during each of 5 independent observation periods. Suppose the probability of an appearance in each period is 0.2. Let \(X\) count the periods with an appearance. Describe the shape of the distribution.
State. Under the stated model, \(X\sim B(5,0.2)\), where success means an appearance in one observation period.
Plan. The two outcomes are appearance and no appearance. There are 5 fixed observation periods; the model assumes the periods are independent and that the appearance probability stays at 0.2. To describe shape, examine the probabilities for all possible counts, 0 through 5.
Do. Apply the binomial probability formula to each possible value:
| Appearances \(k\) | \(P(X=k)\) |
|---|---|
| 0 | 0.3277 |
| 1 | 0.4096 |
| 2 | 0.2048 |
| 3 | 0.0512 |
| 4 | 0.0064 |
| 5 | 0.0003 |
For example, \(P(X=2)=\binom{5}{2}(0.2)^2(0.8)^3=10(0.04)(0.512)=0.2048\). The displayed probabilities sum to 1.0000. The largest probabilities are at 0 and 1 appearances, while the probabilities taper off across the higher counts.
Conclude. The distribution is right-skewed: probability is concentrated at the lower counts, with a tail extending toward 5 appearances. Its mean is \(np=5(0.2)=1\), which helps locate the distribution’s center. The shape description comes from the pattern of probabilities, not from the mean alone.
Worked Example: How Increasing \(n\) Changes the Pattern
Worked Example: How Increasing \(n\) Changes the Pattern
Compare the wildlife-camera model above, \(X\sim B(5,0.2)\), with a model that has 10 independent observation periods and the same appearance probability, 0.2. Let \(Y\) count appearances in the 10-period model. How does the shape compare?
State. The second model is \(Y\sim B(10,0.2)\). Success remains an appearance in one period.
Plan. The model assumes binary outcomes, a fixed 10 periods, independent periods, and the same success probability of 0.2. Compare its probability pattern with the earlier \(B(5,0.2)\) table. The value of \(p\) is unchanged, while \(n\) doubles.
Do. Calculate the probabilities for \(Y\) from 0 to 10. The binomial formula is \(P(Y=k)=\binom{10}{k}(0.2)^k(0.8)^{10-k}\).
| Appearances \(k\) | \(P(Y=k)\), rounded |
|---|---|
| 0 | 0.1074 |
| 1 | 0.2684 |
| 2 | 0.3020 |
| 3 | 0.2013 |
| 4 | 0.0881 |
| 5 | 0.0264 |
| 6 | 0.0055 |
| 7 | 0.0008 |
| 8 | 0.0001 |
| 9 | 0.0000 |
| 10 | 0.0000 |
For instance, \(P(Y=2)=\binom{10}{2}(0.2)^2(0.8)^8=45(0.04)(0.16777216)=0.301989888\), or 0.3020 rounded to four decimal places. The displayed rounded probabilities total 1.0000. The mean changes from \(5(0.2)=1\) in the first model to \(10(0.2)=2\) in the second. The standard deviations are \(\sqrt{5(0.2)(0.8)}=\sqrt{0.8}\approx0.8944\) and \(\sqrt{10(0.2)(0.8)}=\sqrt{1.6}\approx1.2649\).
Conclude. Both distributions are right-skewed because \(p=0.2\), but the 10-period distribution has its probability concentrated around a larger count and spreads across more possible counts. Its shape is less concentrated at the very lowest counts relative to its center, though it still has a tail toward higher counts. Doubling \(n\) does not make the distribution symmetric; it changes the location and spread as well as the visual pattern.
Worked Example: A Left-Skewed Distribution
Worked Example: A Left-Skewed Distribution
Suppose a device has probability 0.8 of passing a check in each of 5 independent tests. Let \(Z\) count the tests passed. Describe the shape and compare it with the \(B(5,0.2)\) wildlife-camera distribution.
State. The device model is \(Z\sim B(5,0.8)\). Success is passing one test, and \(Z\) counts passes.
Plan. Each test has two outcomes, there are 5 fixed tests, and the model assumes independent tests with the same pass probability of 0.8. Because \(0.8=1-0.2\), compare the probabilities with those for a \(B(5,0.2)\) count: each count of passes corresponds to the same probability as the complementary count of failures.
Do. The probabilities are the reverse of those in the earlier \(B(5,0.2)\) table:
| Passes \(k\) | \(P(Z=k)\) |
|---|---|
| 0 | 0.0003 |
| 1 | 0.0064 |
| 2 | 0.0512 |
| 3 | 0.2048 |
| 4 | 0.4096 |
| 5 | 0.3277 |
For example, \(P(Z=4)=\binom{5}{4}(0.8)^4(0.2)^1=5(0.4096)(0.2)=0.4096\). The displayed probabilities total 1.0000. The mean is \(np=5(0.8)=4\).
Conclude. The distribution is left-skewed: most of the probability is at 4 or 5 passes, and the tail extends toward the lower counts. Compared with \(B(5,0.2)\), the pattern is its mirror image: a count of \(k\) passes has the same probability as \(5-k\) failures in the other model.
Common Mistakes and AP Exam Tips
- Confusing the direction of skew. A small \(p\) means lower success counts are more common, with the tail extending toward higher counts: right-skewed. A large \(p\) reverses this pattern: left-skewed.
- Calling a distribution symmetric just because it has a mean. Every binomial distribution has a mean, but only the equal-success-and-failure case \(p=0.5\) gives the binomial symmetry described here.
- Assuming the tallest bar must be at the mean. The mean is a probability-weighted center. It need not be an integer or coincide with the most likely count. Describe shape from the probabilities or histogram.
- Ignoring \(n\). Keeping \(p\) fixed does not keep the histogram unchanged. A different \(n\) changes the range of possible counts, the mean, and the standard deviation.
- Using “skewed” without naming the tail’s direction. State which side contains most of the probability and which way the tail extends. For example, “The distribution is right-skewed, with most probability at lower appearance counts and a tail toward higher counts.”
- Treating a visual pattern as a condition check. A histogram that looks binomial does not establish independence or constant probability. Check the BINS conditions separately, as in the earlier tutorials.
For full-credit communication, name the context-specific count, identify \(n\) and \(p\), and describe where the probability is concentrated and the direction of the tail. If comparing two distributions, state what changed and support the shape comparison with their probabilities or histogram patterns.
Key Takeaway
The shape of a binomial distribution depends on both \(n\) and \(p\). A probability of success equal to 0.5 produces symmetry; probabilities below or above 0.5 tend to produce right- or left-skew, respectively. Use the pattern of probabilities to support your description, and keep shape distinct from center and spread.
Check Your Understanding
For each question, describe the pattern in context and use the binomial parameters to support your answer.
- A count follows \(X\sim B(8,0.5)\). What shape do you expect, and what feature of \(p\) supports that description?
- A community garden models the number of seeds that sprout in 6 independent pots, with probability 0.1 of sprouting in each pot. Which way does the distribution tend to be skewed? Describe where the probability is concentrated and the tail’s direction.
- A second device model has \(n=12\) and \(p=0.9\), where success is passing a test. Describe the expected direction of skew in the count of passes.
- Two binomial models have \(p=0.3\), but one has \(n=5\) and the other \(n=20\). Name one way their histograms must differ and one shape feature they tend to share.
- Why is the mean alone not enough to decide whether a binomial distribution is symmetric or skewed?