Using Standard Deviations to Describe a Mound-Shaped Distribution
In Z-Scores: Standardizing a Value, you learned that a z-score measures how many standard deviations a value is above or below the mean. When a distribution is approximately bell-shaped, that distance also helps describe where observations tend to fall. The empirical rule gives useful approximate percentages for this kind of distribution.
The rule is useful for questions such as “About what percent of values fall between these two measurements?” or “Roughly how many observations are within two standard deviations of the mean?” It does not require the individual data values, but it does require a distribution whose shape is reasonably close to a symmetric, single-peaked mound. As in Describing Shape: Symmetric, Skewed, Uniform and Unimodal, Bimodal, and Multimodal Distributions, use the overall pattern—not one uneven bar—to judge the shape.
The Percentages Around the Mean
The three headline percentages describe regions centered at the mean. If a distribution has mean \(\mu\) and standard deviation \(\sigma\), the regions extend from \(\mu-\sigma\) to \(\mu+\sigma\), from \(\mu-2\sigma\) to \(\mu+2\sigma\), and from \(\mu-3\sigma\) to \(\mu+3\sigma\). For sample data, the corresponding summaries are \(\bar{x}\) and \(s\).
Because a bell-shaped distribution is approximately symmetric, each percentage is divided approximately equally between the two sides of the mean. About 34% lies between the mean and one standard deviation above it, and about 34% lies between the mean and one standard deviation below it. The 68% total is therefore \(34\%+34\%\).
The regions farther from the mean can also be found by subtraction. About 95% is within two standard deviations, while about 68% is within one. The difference, \(95\%-68\%=27\%\), lies between one and two standard deviations from the mean on both sides. Split evenly, that is about 13.5% in each band. Similarly, about \(99.7\%-95\%=4.7\%\) lies between two and three standard deviations from the mean, or about 2.35% in each band.
The remaining \(100\%-99.7\%=0.3\%\) lies more than three standard deviations from the mean. For a roughly symmetric mound, about 0.15% lies in each tail. These smaller pieces are useful when a question asks for a percentage in only part of the distribution.
- From the mean to 1 standard deviation in either direction: about 34% per side.
- Between 1 and 2 standard deviations in either direction: about 13.5% per side.
- Between 2 and 3 standard deviations in either direction: about 2.35% per side.
- Beyond 3 standard deviations: about 0.15% in each tail.
A Reliable Method for Percent-Between Questions
A “percent between” question may ask for a region that is not one of the three centered intervals. The key is to locate both endpoints relative to the mean in standard-deviation units, mark the relevant regions, and add the pieces between the endpoints. When a requested region crosses the mean, include the appropriate portion on each side. When it does not cross the mean, do not accidentally include the other side of the distribution.
Confirm that the distribution is approximately symmetric and mound-shaped, with one main peak and no strong skew or unusual feature that makes the rule inappropriate.
Express each value as the mean plus or minus a number of standard deviations. The z-score idea from the earlier tutorial can help locate a value.
Use the 68%, 95%, and 99.7% totals—or the smaller pieces derived from them—to identify just the area between the endpoints.
State the approximate percent in context. If asked for a count, multiply the estimated proportion by the total number of observations and describe the result as approximate.
A useful sketch can prevent sign and region errors. Put the mean in the middle, mark one, two, and three standard deviations on both sides, and label each band with its approximate percent. The sketch does not need a vertical scale; it is a map for deciding which percentages to add.
Worked Examples: Applying the Empirical Rule
Worked Example: Percent Between One Standard Deviation Below and Two Above
A fictional distribution of daily water use per household is approximately bell-shaped, with mean 240 liters and standard deviation 30 liters. About what percent of households use between 210 and 300 liters per day?
Check and locate. The values are one standard deviation below the mean and two standard deviations above it: \(210=240-30\), and \(300=240+2(30)\). The stated distribution is approximately bell-shaped, so applying the empirical rule is reasonable for an estimate.
Mark the regions. The interval from 210 to 300 includes the band from one standard deviation below the mean up to the mean, about 34%. Above the mean, it includes the band from the mean to one standard deviation above, about 34%, and the band from one to two standard deviations above, about 13.5%.
Check the total another way. The full interval within two standard deviations is about 95%. The requested interval excludes the band below \(-1\sigma\), about 16% altogether from \(-2\sigma\) to \(-1\sigma\) and beyond \(-2\sigma\) only as relevant? This subtraction would not isolate the requested region cleanly, so use the marked bands: from \(-1\sigma\) to \(+2\sigma\) is 34%+34%+13.5%=81.5%.
Conclude in context. About 81.5% of households in this fictional distribution use between 210 and 300 liters of water per day. This is an estimate based on the empirical rule, not a claim that exactly 81.5% of the households do so.
Worked Example: A Range Within Two Standard Deviations
A fictional set of repair times for a certain type of appliance is approximately bell-shaped, with mean 52 minutes and standard deviation 6 minutes. Find the interval containing about 95% of repair times.
Plan. The empirical rule places about 95% of observations within two standard deviations of the mean. The interval is therefore the mean minus two standard deviations through the mean plus two standard deviations.
Calculate the endpoints. Two standard deviations equal \(2(6)=12\) minutes. The lower endpoint is \(52-12=40\) minutes, and the upper endpoint is \(52+12=64\) minutes.
Check. The midpoint of 40 and 64 is \((40+64)/2=52\), the stated mean. Each endpoint is 12 minutes from the mean, and \(12/6=2\) standard deviations.
Conclude in context. About 95% of repair times in this fictional, approximately bell-shaped distribution are between 40 and 64 minutes. The empirical rule gives an approximate interval; it does not guarantee that exactly 95% of observed repair times fall there.
Worked Example: Estimating a Number of Observations in a Band
A fictional group of 200 seedlings has approximately bell-shaped heights, with mean 18 centimeters and standard deviation 2 centimeters. Estimate how many seedlings are between 20 and 24 centimeters tall.
State and plan. The interval from 20 to 24 centimeters runs from one to three standard deviations above the mean: \(20=18+1(2)\), and \(24=18+3(2)\). For an approximately bell-shaped distribution, the band from one to two standard deviations above the mean contains about 13.5%, and the band from two to three above contains about 2.35%. We will add these percentages and multiply by 200.
Do. The estimated percentage is:
As a proportion, \(15.85\%=0.1585\). The estimated number of seedlings is:
Check. The two bands contain about 13.5% and 2.35%, so their combined percentage is 15.85%. Also, \(31.7/200=0.1585\), which is 15.85% of the group. Since a count of seedlings must be a whole number, the estimate is about 32 seedlings.
Conclude in context. About 32 of the 200 fictional seedlings are estimated to be between 20 and 24 centimeters tall. The estimate need not match the actual count exactly because both the empirical-rule percentage and the resulting count are approximate.
Worked Example: Percent Outside Two Standard Deviations
A fictional distribution of package weights is approximately bell-shaped. About what percent of packages weigh either less than two standard deviations below the mean or more than two standard deviations above it?
Use the centered percentage. About 95% of observations lie within two standard deviations of the mean. Therefore, the percent outside that interval is the remaining percentage:
Use symmetry as a check. A roughly symmetric distribution divides that 5% approximately equally between its two tails: \(5\%/2=2.5\%\) in each tail. Adding the tails gives \(2.5\%+2.5\%=5\%\).
Conclude in context. About 5% of package weights are estimated to be more than two standard deviations from the mean, counting both low and high weights. About 2.5% are estimated in each tail. This is an approximate description of the fictional mound-shaped distribution.
When the Rule Is Appropriate—and What It Does Not Say
The empirical rule is a rule of thumb for a distribution that is approximately bell-shaped: roughly symmetric, with one main mound and no severe departure from that pattern. A histogram is useful for judging whether this description fits. The shape matters because the percentages rely on the balanced pattern of a mound; the same percentages should not be assumed for a strongly skewed, bimodal, or irregular distribution.
The rule does not say that every bell-shaped data set has exactly 68%, 95%, and 99.7% in the stated intervals. Those are approximate benchmarks. Actual sample percentages can differ, especially in smaller data sets, and a few observations may fall beyond three standard deviations. Use words such as “about” or “approximately” when reporting the result.
The empirical rule describes proportions of observations in regions around the mean. It does not identify which particular observations are in those regions, and it does not provide an exact value for a percentile or probability beyond the approximations it states. If the distribution is not reasonably mound-shaped, use the graph or other information available rather than forcing the empirical rule onto it.
Common Mistakes and AP Exam Tips
- Applying the rule to any distribution. A mean and standard deviation alone do not justify the 68-95-99.7 percentages. A full-credit response first notes that the distribution is approximately bell-shaped or explains why the rule is appropriate.
- Using 68%, 95%, or 99.7% for the wrong interval. Those percentages describe regions centered at the mean. For example, 95% is within two standard deviations on both sides—not between one and two standard deviations on one side.
- Forgetting to split a two-sided band. About 27% lies between one and two standard deviations on both sides combined, so one side has about \(27\%/2=13.5\%\), assuming approximate symmetry.
- Adding regions outside the requested interval. Mark the endpoints first. For a percent-between question, add only the bands that actually lie between them.
- Reporting an estimate as exact. Say “about 81.5%” or “approximately 32 seedlings,” not that exactly that proportion or count must occur.
- Giving a count when asked for a percent, or vice versa. To convert an estimated percentage to a count, change the percent to a proportion and multiply by the number of observations. Include the count’s context and acknowledge rounding.
A strong AP response identifies why the rule is suitable, shows which bands are being combined, and gives the estimate in context. For example: “Because the distribution is approximately bell-shaped, the interval from one standard deviation below the mean to two above it contains about \(34\%+34\%+13.5\%=81.5\%\) of observations.” This makes the reasoning visible and keeps the answer appropriately approximate.
Check Your Understanding
Assume each distribution below is approximately bell-shaped. Use the empirical rule and show which regions you combine.
- A distribution has mean 70 and standard deviation 5. About what percent of observations are between 65 and 75?
- For the same distribution, about what percent of observations are between 70 and 80?
- A fictional set of 400 measurements is mound-shaped. About how many measurements are within three standard deviations of the mean?
- For a mound-shaped distribution, about what percent lies between two and three standard deviations above the mean?
- Why should you not use the empirical rule automatically for a strongly right-skewed distribution?