What Makes a Distribution Normal?
A normal distribution is a continuous probability distribution with a smooth, symmetric, bell-shaped curve. Many measurements, including some test scores, can be modeled approximately by a normal distribution. Its shape gives us a useful way to describe where values tend to fall and how much of the distribution lies in different intervals.
In Normal Approximation to the Binomial, a normal curve was used to approximate probabilities for certain binomial counts. Here, we focus on the features of the curve itself: how symmetry and the mean locate its center, how standard deviation describes its spread, and how the empirical rule estimates proportions in intervals around the center.
“Symmetric” means the left and right sides are mirror images around the center. In a normal distribution, equal-width intervals equally far above and below the mean have equal probabilities. The curve has one peak at the mean and tapers toward both tails. The tails continue indefinitely, getting closer to the horizontal axis without touching it.
The mean and standard deviation play distinct roles. The mean locates the center: shifting the mean moves the whole curve to a different position on the measurement scale. The standard deviation controls the spread: a larger standard deviation makes the curve wider and flatter, while a smaller standard deviation makes it narrower and taller. The total area under each curve is 1, so a wider curve does not represent more observations overall; it spreads the same total area across a wider range.
Standard deviation is measured in the same units as the observations. For example, if test scores are measured in points, the mean and standard deviation are both in points. As discussed in Interpreting Standard Deviation of a Random Variable, standard deviation describes typical distance from the mean; it is not a rule that every value must be close to the mean.
The Empirical Rule
For a normal distribution, the empirical rule gives approximate proportions within one, two, and three standard deviations of the mean. It is also called the 68–95–99.7 rule. These percentages are approximate, not exact guarantees about any particular group of observations.
The symmetry of the curve lets us split each centered interval evenly. About 34% of values lie between the mean and 1 standard deviation above it, and about 34% lie between the mean and 1 standard deviation below it. About 2.5% lie beyond 2 standard deviations in each tail, since about 5% lie outside the interval from 2 standard deviations below to 2 above. About 0.15% lie beyond 3 standard deviations in each tail, since about 0.3% lie outside the interval from 3 standard deviations below to 3 above.
Subtracting the empirical-rule percentages also estimates the proportions in bands. About \(95\%-68\%=27\%\) of values lie between 1 and 2 standard deviations from the mean, counting both sides together. By symmetry, about half of that, or 13.5%, is on each side. Likewise, about \(99.7\%-95\%=4.7\%\) lies between 2 and 3 standard deviations from the mean, or about 2.35% on each side.
To apply the rule, first use the mean and standard deviation to find the numerical endpoints. Then match the requested interval to one or more empirical-rule regions. A range from the mean to 1 standard deviation above the mean is about 34%; a range from 1 standard deviation below to 2 below is about 13.5%. Keep track of whether a question asks for one tail, both tails, or a band between two cutoffs.
Worked Example: Locating Test Scores Around the Mean
Worked Example: Locating Test Scores Around the Mean
Suppose a school’s model for a particular exam describes scores with a normal distribution having a mean of 72 points and a standard deviation of 8 points. Estimate the percentage of scores from 64 to 80, from 56 to 88, and from 48 to 96.
Identify the intervals. The mean is 72 points. One standard deviation below and above the mean gives \(72-8=64\) and \(72+8=80\). Two standard deviations gives \(72-2(8)=56\) and \(72+2(8)=88\). Three standard deviations gives \(72-3(8)=48\) and \(72+3(8)=96\).
Apply the empirical rule. The interval 64 to 80 is within 1 standard deviation of the mean, so it contains about 68% of scores. The interval 56 to 88 is within 2 standard deviations, so it contains about 95%. The interval 48 to 96 is within 3 standard deviations, so it contains about 99.7%.
Interpret. Under this normal model, about 68% of scores are between 64 and 80 points, about 95% are between 56 and 88 points, and about 99.7% are between 48 and 96 points. These are model-based approximations, not claims that the exact percentages in every group of students must match those values.
Worked Example: Estimating a Band and a Tail
Worked Example: Estimating a Band and a Tail
For the same modeled exam, the mean is 72 points and the standard deviation is 8 points. Estimate the percentage of scores from 80 to 88 and the percentage above 88.
Find the cutoffs in standard deviations. Since \(80=72+8\), a score of 80 is 1 standard deviation above the mean. Since \(88=72+2(8)\), a score of 88 is 2 standard deviations above the mean.
Estimate the band from 80 to 88. The area between 1 and 2 standard deviations above the mean is half of the area between 1 and 2 standard deviations from the mean on either side. The two-sided band is approximately \(95\%-68\%=27\%\). By symmetry, the upper-side band is about \(27\%/2=13.5\%\).
Estimate the percentage above 88. A score of 88 is 2 standard deviations above the mean. About 95% of scores are within 2 standard deviations, leaving about \(100\%-95\%=5\%\) outside that interval. Symmetry splits the remaining area equally between the two tails, so about \(5\%/2=2.5\%\) of scores are above 88.
Interpret. Under the model, about 13.5% of scores fall between 80 and 88 points, and about 2.5% exceed 88 points. The symmetry is essential to both estimates: it allows us to divide a two-sided area evenly between the upper and lower sides.
Worked Example: Compare Spread in Two Test-Score Models
Worked Example: Compare Spread in Two Test-Score Models
Two practice exams are modeled with normal distributions. Scores on Exam A have a mean of 70 points and a standard deviation of 5 points. Scores on Exam B have a mean of 70 points and a standard deviation of 10 points. Compare the approximate percentages of scores within 5 points of the mean for each exam.
Exam A. Five points is exactly 1 standard deviation for Exam A. The interval within 5 points of the mean is \(70-5=65\) to \(70+5=75\). By the empirical rule, about 68% of Exam A scores fall in this interval.
Exam B. Five points is half of Exam B’s standard deviation. The requested interval is still 65 to 75, but it is not an interval of 1, 2, or 3 standard deviations for this model. The empirical rule alone does not give an estimate for exactly half a standard deviation on either side, so we should not claim that the percentage is 68%. We can still compare the curves: Exam B has the same center but a larger standard deviation, so its scores are more spread out. Therefore, a smaller proportion of Exam B scores than Exam A scores will be within 5 points of the mean.
Interpret. The means are equal, so both models are centered at 70 points. Exam A’s smaller standard deviation means its scores are more concentrated near 70. The empirical rule estimates about 68% within 5 points for Exam A; it does not, by itself, provide an exact percentage for Exam B’s half-standard-deviation interval. This distinction avoids treating the 68% rule as if it applied to any interval chosen around the mean.
When the Empirical Rule Is Appropriate
The empirical rule applies to a normal distribution or to data whose distribution is reasonably modeled by a normal curve. It is not a universal percentage rule for every data set. A distribution that is strongly skewed, has multiple peaks, or has an unusual shape may not have the approximate 68%, 95%, and 99.7% proportions.
A problem may explicitly state that a distribution is normal. If it instead gives a set of observed scores, consider whether the shape is approximately symmetric, single-peaked, and bell-shaped before using the rule. A few unusual observations do not automatically invalidate a normal model, but a substantial skew or a second cluster of values is a warning that the model may not describe the data well.
The empirical rule gives approximate proportions, not individual predictions. If a score is more than 2 standard deviations from the mean, it is in one of the relatively small tails, but that does not make the score impossible. Even scores more than 3 standard deviations from the mean can occur; the rule estimates that about 0.3% of values are outside that central interval.
Common Mistakes and AP Exam Tips
- Confusing the mean with the spread. The mean identifies the center; standard deviation describes spread. If two distributions share a mean but have different standard deviations, they have the same center, not the same shape or concentration.
- Using 68% for any interval around the mean. The 68% estimate is for the interval within 1 standard deviation on both sides. From the mean to 1 standard deviation above it is only half of that area, about 34%.
- Forgetting to split a two-tail percentage. About 5% lies outside 2 standard deviations in total, but symmetry puts about 2.5% in each tail. State whether the requested area is one tail or both.
- Subtracting percentages without identifying the region. The difference \(95\%-68\%=27\%\) covers both bands between 1 and 2 standard deviations. If the question asks for just the upper band, divide by 2 to get 13.5%.
- Treating approximations as exact counts. “About 68%” is a proportion estimate. To estimate a count from a specified group size, multiply by that size and remember the result is approximate; it need not be a whole number before rounding.
- Applying the rule to an unsuitable shape. The percentages are not guaranteed for skewed or multimodal distributions. A full-credit answer connects use of the empirical rule to a stated normal model or a distribution that is reasonably bell-shaped.
A strong AP response identifies the mean and standard deviation, translates the requested score cutoffs into distances from the mean, and explains which empirical-rule area applies. If symmetry is used to divide an area, say so. Finish by interpreting the approximate percentage in the test-score context rather than leaving a bare number.
Key Takeaway
The normal distribution’s symmetry places its mean at the center, while its standard deviation determines how concentrated or spread out values are around that center. The empirical rule turns those features into useful approximate percentages for intervals one, two, and three standard deviations from the mean.
Check Your Understanding
Use the empirical rule and the roles of the mean and standard deviation to answer each question.
- A test-score model is normal with a mean of 84 points and a standard deviation of 6 points. Find the interval that contains about 68% of scores.
- For the model in question 1, what percentage of scores are estimated to fall between 72 and 78 points? Show how you identify the relevant band.
- For a normal distribution, about what percentage of values lie above 2 standard deviations greater than the mean? Explain how symmetry helps.
- Two normal score distributions have the same mean, but one has a larger standard deviation. Which distribution is more spread out, and what does that imply about the proportion close to the mean?
- A data set is strongly right-skewed. Is it appropriate to assume that about 95% of its values lie within 2 standard deviations of its mean? Explain.