From the Empirical Rule to an Area Estimate
In Features of a Normal Distribution, you learned the 68-95-99.7 rule: for a normal distribution, about 68% of values are within 1 standard deviation of the mean, about 95% are within 2 standard deviations, and about 99.7% are within 3 standard deviations. This tutorial uses those benchmarks to estimate the proportion in a particular interval or tail.
The main technique is to mark the mean and the values 1, 2, and 3 standard deviations on each side of it, then divide the curve into bands. Use symmetry to assign areas to the bands, add the bands that match the requested region, and check that the shaded region on a sketch agrees with your answer.
The rule gives cumulative areas around the center, not the area of every individual band. For example, the area between 1 and 2 standard deviations above the mean is found by taking the area within 2 standard deviations, subtracting the area within 1 standard deviation, and splitting what remains equally between the two sides. A band ledger makes that reasoning easier to organize.
| Region on one side of the mean | Approximate area | Reason |
|---|---|---|
| Mean to 1 standard deviation | 0.34 | Half of the 0.68 central area |
| Between 1 and 2 standard deviations | 0.135 | \((0.95-0.68)/2=0.135\) |
| Between 2 and 3 standard deviations | 0.0235 | \((0.997-0.95)/2=0.0235\) |
| Beyond 3 standard deviations | 0.0015 | \((1-0.997)/2=0.0015\) |
The corresponding two-sided areas are 0.68 within 1 standard deviation, 0.95 within 2, and 0.997 within 3. For instance, the area above 2 standard deviations is the area from 2 to 3 plus the area beyond 3: \(0.0235+0.0015=0.025\). That is also half of the 0.05 area outside 2 standard deviations.
Locate the Boundaries, Then Sketch the Region
When the problem gives raw values rather than standard-deviation marks, use the z-score formula from Calculating a z-Score to locate each value:
A z-score tells you how many standard deviations a value lies from the mean. A value with \(z=-2\) is 2 standard deviations below the mean, while a value with \(z=1\) is 1 standard deviation above it. Once the boundaries are located, sketch a bell-shaped curve centered at the mean, mark the boundaries, and shade the requested region.
A sketch is a reasoning check, not a source of exact measurements. It helps catch mistakes such as adding a left-tail area when the question asks for the area to the right, or including a band outside the interval. Because the normal curve is symmetric, matching bands on opposite sides have equal areas. The curve also has two tails, so an area outside a central interval must be split between them when the question asks for only one tail.
Use the empirical rule when the variable’s distribution is normal or approximately normal. The rule is not a guarantee for a strongly skewed or irregular distribution.
Use the mean and standard deviation, or calculate z-scores to express raw-value boundaries in standard-deviation units.
Draw the bell-shaped curve, label the center and relevant standard-deviation marks, and shade exactly the interval or tail in the question.
Use the empirical-rule benchmarks and symmetry to find the approximate area of the shaded region.
Confirm that the result is between 0 and 1 and that its size makes sense for the sketch. State what the approximate proportion means in context.
Worked Example: An Interval from Two Standard Deviations Below to One Above
Worked Example: An Interval from Two Standard Deviations Below to One Above
A fictional bottling line models the amount of juice in a bottle as approximately normal, with mean 500 milliliters and standard deviation 12 milliliters. Estimate the proportion of bottles containing between 476 and 512 milliliters.
State. Let \(X\) be the amount of juice in a bottle, in milliliters. The model is approximately normal with \(\mu=500\) and \(\sigma=12\). We want the proportion for which \(476\leq X\leq512\).
Plan. The stated model is approximately normal, so the empirical rule is appropriate for an estimate. Convert the two endpoints to z-scores, then use a sketch and add the areas between those marks.
Do. Calculate the z-score of each boundary:
The requested interval runs from \(z=-2\) to \(z=1\). On a sketch, shade from 2 standard deviations below the mean to 1 standard deviation above it. The region includes the band from \(-2\) to \(-1\), whose area is about 0.135, the band from \(-1\) to the mean, whose area is about 0.34, and the band from the mean to \(1\), whose area is about 0.34.
Conclude. About 0.815, or 81.5%, of the bottles are estimated to contain between 476 and 512 milliliters under this model. The sketch check is sensible: the interval includes most of the curve, but it leaves out the portion below \(-2\) and the portion above \(1\).
Worked Example: Estimating a Right-Tail Proportion
Worked Example: Estimating a Right-Tail Proportion
In a fictional city, daily bicycle trips on a popular route are modeled as approximately normal with mean 24 minutes and standard deviation 4 minutes. Estimate the proportion of trips that take more than 32 minutes.
Locate the cutoff. For a trip time of 32 minutes:
Thus, the question asks for the area to the right of \(z=2\). On a sketch, mark the mean at 24 minutes, mark 32 minutes two standard deviations to its right, and shade only the right tail beyond 32.
The area outside 2 standard deviations is approximately \(1-0.95=0.05\), but that is the combined area in both tails. By symmetry, the right tail is half of that:
About 0.025, or 2.5%, of the trips are estimated to take more than 32 minutes. If the model is used for 1,000 trips, the estimated number taking more than 32 minutes is \(1000(0.025)=25\). This is an expected count based on the estimate, not a promise that exactly 25 trips will exceed 32 minutes.
The sketch helps identify an easy-to-miss issue: 0.05 is the area in both tails together, not the area in the right tail alone. Since the requested region is just one tail and the curve is symmetric, the answer is about half of 0.05.
Worked Example: Adding Two Bands Between One and Three Standard Deviations
Worked Example: Adding Two Bands Between One and Three Standard Deviations
A fictional greenhouse models the length of a certain type of seedling as approximately normal with mean 50 centimeters and standard deviation 5 centimeters. Estimate the proportion of seedlings with lengths between 55 and 65 centimeters.
First locate the two endpoints. For 55 centimeters, \(z=(55-50)/5=1\). For 65 centimeters, \(z=(65-50)/5=3\). The interval is therefore from 1 to 3 standard deviations above the mean. A sketch should show two bands on the right side: the band from 1 to 2 and the band from 2 to 3.
The first band has approximate area \((0.95-0.68)/2=0.135\). The second has approximate area \((0.997-0.95)/2=0.0235\). Add the areas because both bands are included:
About 15.9% of the seedlings are estimated to have lengths from 55 to 65 centimeters. For a group of 2,000 seedlings modeled this way, the corresponding estimated count is \(2000(0.1585)=317\) seedlings.
Check the sketch before accepting the result. The interval is only one side of the curve, so its area should be smaller than the combined area within 2 standard deviations. The estimate, about 15.9%, is indeed smaller than 95%. It is also larger than the area in the 2-to-3 band alone because the interval includes the entire 1-to-2 band as well.
Why These Estimates Are Approximate
The values 68%, 95%, and 99.7% are rounded benchmarks for a normal distribution. Areas calculated from them inherit that approximation. For example, 0.135 is the band estimate obtained from the rounded 68% and 95% benchmarks; it is not an exact area for every normal distribution. Report a result as approximate and avoid implying more precision than the rule supports.
The empirical rule is useful when a quick estimate is enough or when the requested endpoints fall at convenient standard-deviation marks. If a problem asks for a more precise normal probability, a normal probability calculation using technology is more appropriate. The empirical rule and a calculator answer need not match exactly because the rule uses rounded benchmark areas.
Also distinguish an estimated proportion from an observed count. If an estimated proportion is 0.025, that means about 2.5% under the model. Multiplying by a group size gives an estimated count, but actual observations can differ. The model describes a distribution; it does not force a particular sample or group to match its expected proportions exactly.
Common Mistakes and AP Exam Tips
- Using a cumulative area as a single-band area. The 95% benchmark is the area within 2 standard deviations on both sides. To get the area between 1 and 2 on one side, subtract 0.68 from 0.95 and divide by 2.
- Forgetting to split a two-sided tail area. The 5% outside 2 standard deviations is split between two tails. A question asking for only the right tail calls for about \(0.05/2=0.025\), not 0.05.
- Shading the wrong region. Before adding areas, mark both endpoints and shade only the values described. A sketch can expose a sign or tail error even when the arithmetic is correct.
- Adding bands that are not in the interval. For an interval from \(z=1\) to \(z=3\), add the 1-to-2 and 2-to-3 bands. Do not add the central region or the tail beyond 3.
- Applying the rule to any distribution. The empirical rule is for normal or approximately normal distributions. It is not justified just because a problem gives a mean and standard deviation.
- Reporting a rule-based estimate as exact. Use words such as “about” or “approximately,” and round sensibly. A result such as 0.1585 should be reported as about 0.159, or 15.9%.
- Giving a number without context. A full-credit response identifies the variable and the event, reports the approximate proportion, and explains what that proportion means for the stated situation.
For a clear AP response, show how the boundaries relate to the mean and standard deviation, identify the shaded part of the curve, show the band areas being combined, and finish with an interpretation in context. If you estimate a count, show the multiplication by the group size and call the result an estimate.
Key Takeaway
The empirical rule turns a few normal-curve benchmarks into quick estimates for many intervals and tails. Translate raw boundaries into standard-deviation marks, sketch and shade the requested region, and add only the band areas it contains. Use symmetry carefully, check whether an outside area includes one tail or both, and explain that the result is approximate.
Check Your Understanding
Use the empirical rule and a sketch to estimate each requested proportion. Show the bands or tail areas you use, and interpret each result in context.
- A normal model has mean 80 and standard deviation 6. Estimate the proportion of values between 74 and 86.
- A normal model has mean 120 and standard deviation 10. Estimate the proportion of values greater than 140.
- A normal model has mean 30 and standard deviation 3. Estimate the proportion of values between 33 and 39.
- Explain why the estimated area below \(z=-2\) is about 0.025 rather than 0.05.
- If approximately 0.34 of a modeled group lies between the mean and 1 standard deviation above the mean, how many people would that estimate for a group of 500? Explain why the actual count need not equal your estimate.