Start with Shape, Not Just a Formula
In Interpreting Simulation Results Against a Model, you compared simulated results with what a stated probability model predicts. Here the question comes earlier: do the data’s shape and summary statistics make a normal model reasonable in the first place? This matters because a normal model is useful only when its shape is a credible representation of the variable being studied.
As covered in Features of a Normal Distribution, a normal distribution is symmetric, single-peaked, and bell-shaped. Its mean is at the center, and its two sides mirror one another. Real data do not need to match a perfect curve, but a noticeably long tail or strong asymmetry can make a normal model a poor choice.
The word “tend” is important. The relationship between mean and median is a clue, not a rule that classifies every data set. A few unusual values can pull the mean, and different shapes can produce similar means and medians. Use several pieces of evidence together rather than relying on one statistic.
Use Several Summary Statistics as Clues
The mean uses every observation, so unusually large or small values can affect it substantially. The median depends on the middle position in the ordered data and is less affected by extreme values. When a few high observations stretch the right tail, the mean is often pulled above the median. When a few low observations stretch the left tail, the mean is often pulled below it.
Quartiles add another useful comparison. The first quartile, \(Q_1\), marks the approximate 25th percentile; the third quartile, \(Q_3\), marks the approximate 75th percentile. Compare the distance from \(Q_1\) to the median with the distance from the median to \(Q_3\). If one side is much longer, that can point to asymmetry. Also note unusually distant minimum or maximum values: a long stretch toward the largest values is a clue of right skew.
A standard deviation describes spread around the mean, but it does not tell you whether the data are symmetric. Two distributions can have similar means and standard deviations while having very different shapes. This is why, in Checking Whether Data Are Approximately Normal, histograms, empirical-rule comparisons, and normal probability plots were treated as complementary evidence. Use summary statistics to guide your judgment, and use a graph when the data are available.
Context matters, too. A group of heights drawn from people of similar ages and other relevant characteristics might have a roughly symmetric, single-peaked distribution. Combining distinct groups can change the shape. Income distributions often have a practical lower limit but can extend far upward, because a small number of people may have much higher incomes than most. That pattern can produce a long right tail and a mean above the median. These are expectations to investigate, not assumptions to impose on every data set.
Worked Example: A Sample of Heights
Worked Example: A Sample of Heights
A fictional school project records the heights, in centimeters, of 10 students in a similar age group. The ordered values are 158, 160, 162, 163, 164, 165, 166, 167, 169, and 171. Judge whether a normal model seems reasonable based on the summaries.
Do. The sum is \(1645\), so the sample mean is \(1645/10=164.5\) cm. With an even number of values, the median is the average of the fifth and sixth observations: \((164+165)/2=164.5\) cm. The mean and median are equal in this sample.
The lower half is 158, 160, 162, 163, 164, so \(Q_1=162\) cm. The upper half is 165, 166, 167, 169, 171, so \(Q_3=167\) cm. The two distances from the median are \(164.5-162=2.5\) cm and \(167-164.5=2.5\) cm. The minimum and maximum are also equally distant from the center: \(164.5-158=6.5\) cm and \(171-164.5=6.5\) cm.
For additional context, the sum of the squared deviations from the mean is \(142\) square centimeters. Thus the sample standard deviation is \(s=\sqrt{142/(10-1)}\approx3.97\) cm, or about \(4.0\) cm. This describes spread but does not, by itself, establish the shape.
Conclude. The matching mean and median, balanced quartile distances, and balanced extremes are consistent with a roughly symmetric distribution. These summaries support considering a normal model for this group of heights, but they do not establish that the population is normal. The sample is small, so a graph and the way the students were selected would also matter.
Worked Example: A Right-Skewed Income Sample
Worked Example: A Right-Skewed Income Sample
A fictional community survey records annual incomes, in thousands of dollars, for 10 adults: 28, 30, 32, 34, 35, 36, 38, 40, 45, and 122. Use the summaries to judge whether a normal model seems reasonable for these observations.
Do. The total is \(440\) thousand dollars, so the mean is \(440/10=44\) thousand dollars. The median is \((35+36)/2=35.5\) thousand dollars. The mean is \(44-35.5=8.5\) thousand dollars above the median, suggesting that larger values are pulling the mean upward.
The lower half is 28, 30, 32, 34, 35, so \(Q_1=32\). The upper half is 36, 38, 40, 45, 122, so \(Q_3=40\). The distance from \(Q_1\) to the median is \(35.5-32=3.5\) thousand dollars, while the distance from the median to \(Q_3\) is \(40-35.5=4.5\) thousand dollars. More strikingly, the largest value, 122, is far above the other observations.
The interquartile range is \(40-32=8\) thousand dollars. As an additional check on the unusually high value, the upper fence using the 1.5-IQR rule is \(40+1.5(8)=52\) thousand dollars. Since \(122>52\), the largest observation is beyond that fence. This supports describing the sample as having an extreme high value and a long right tail.
Conclude. The mean being above the median and the extreme high income both indicate right skew. A symmetric normal model would not represent this sample’s shape well, so it is not a reasonable description of these observations. The conclusion is about this sample and the evidence provided; it does not claim that every group of incomes has exactly the same shape.
Worked Example: Equal Mean and Median, but Not a Normal Shape
Worked Example: Equal Mean and Median, but Not a Normal Shape
Imagine a fictional collection of heights, in centimeters, that combines two distinct groups. Its 12 ordered values are 150, 151, 152, 153, 154, 155, 175, 176, 177, 178, 179, and 180. Could the equal mean and median make a normal model reasonable?
Do. The first six values sum to \(915\), and the last six sum to \(1065\), for a total of \(1980\). The mean is \(1980/12=165\) cm. The median is the average of the sixth and seventh values, \((155+175)/2=165\) cm. Thus the mean equals the median.
The first quartile is the average of the third and fourth values in the lower half: \(Q_1=(152+153)/2=152.5\) cm. The third quartile is the average of the third and fourth values in the upper half: \(Q_3=(177+178)/2=177.5\) cm. These quartiles are equally spaced around the median, but the ordered values show a large gap between 155 and 175, with no observations near the center.
Conclude. Equal mean and median do not make a normal model reasonable here. The values form two separated clusters rather than one central peak. A normal distribution is single-peaked, so it would fail to represent this important feature. The summaries alone could miss the gap; inspecting the values or a graph is essential.
Make a Judgment, Not a Proof
These examples show why a normal-model judgment should draw on several clues. A mean noticeably above the median, uneven quartile spacing, and unusually high values point toward right skew. A mean noticeably below the median and a long stretch toward small values point toward left skew. Similar mean and median, balanced quartile spacing, and no striking extremes are more consistent with symmetry, but they do not guarantee a single-peaked, bell-shaped distribution.
There is no universal cutoff for how far apart the mean and median may be before a normal model is rejected. Their difference must be considered in relation to the variable’s units and spread. For example, a difference of 8.5 thousand dollars in the income example is meaningful alongside the median and the extreme observation; the number is not a general rule for all income data.
A normal model can also be unsuitable for reasons that summaries do not reveal. The data might have two clusters, a gap, or a shape that changes across subgroups. Conversely, a modest amount of irregularity in a sample does not automatically rule out a useful approximation. Use the context and the distribution’s overall shape, not a single summary statistic, to make the decision.
Common Mistakes and AP Exam Tips
- Treating mean and median as a formal test. A mean close to the median is a clue of symmetry, not proof of normality. Full-credit reasoning describes it as evidence and checks other features.
- Ignoring the direction of the tail. In a right-skewed distribution, high values tend to pull the mean above the median. In a left-skewed distribution, low values tend to pull it below. Name the direction and connect it to the context.
- Assuming standard deviation measures shape. Standard deviation measures spread around the mean. It does not tell whether the distribution is symmetric, skewed, or multi-peaked.
- Assuming all measurements are normal. Heights may be roughly symmetric within a suitable group, but mixing different ages or other populations can alter the shape. Describe the group being modeled.
- Calling an unusual value proof of skewness or model failure by itself. An extreme observation is a reason to investigate shape and context. Explain how it relates to the rest of the distribution.
- Using summaries while ignoring visible clusters. Equal mean and median can occur in a distribution that is not single-peaked. Look at the ordered data, histogram, or normal probability plot when available.
Key Takeaway
Summary statistics can reveal clues about a distribution’s symmetry and tails. The mean-median relationship, quartile spacing, and extreme values are useful together, especially when interpreted in context. A normal model is most plausible when the evidence supports a roughly symmetric, single-peaked shape; summaries can support that judgment, but they cannot replace examination of the distribution.
Check Your Understanding
For each situation, explain what the summaries suggest about shape and whether they support considering a normal model.
- A sample of daily wait times has mean 18 minutes, median 12 minutes, and a few especially long waits. What shape is suggested, and why?
- A roughly symmetric sample has \(Q_1=42\), median \(=50\), and \(Q_3=58\). What do the quartile distances suggest? Do they prove the data are normal?
- A distribution has mean 74 and median 81, with some observations far below the rest. What direction of skewness is suggested?
- Why is the standard deviation alone not enough to decide whether a normal model fits?
- A data set has equal mean and median but two clearly separated clusters. Explain why a normal model may still be inappropriate.