Tutorials › AP Statistics › Checking Whether Data Are Approximately Normal

Normal distributions · Tutorial 377 of 1000

Checking Whether Data Are Approximately Normal

Use visual shape, empirical-rule percentages, and normal probability plots to decide whether a normal model is a reasonable approximation for quantitative data.

Intermediate 9 min read

What You'll Learn

  • Describe the features of a histogram that support or weaken a normal model.
  • Use symmetry and the empirical rule as checks, not as automatic proof of normality.
  • Compare observed percentages within one, two, and three standard deviations with the empirical-rule benchmarks.
  • Read a normal probability plot and recognize patterns that may indicate skewness or outliers.
  • Explain why assessing the shape of individual data is different from assessing a sampling distribution.

Why Check the Shape of the Data?

In Normal Models for Sample Means and Proportions, we used normal models to describe sample statistics when the relevant conditions were met. Before treating quantitative data as approximately normal, it helps to examine the data’s shape. A normal model is symmetric and single-peaked, so a histogram can reveal whether that description seems reasonable. A normal probability plot gives another visual check.

These checks are about the distribution represented by the data. They do not prove that a population is exactly normal, and they do not replace the conditions for a normal model of a sampling distribution. For example, as explained in the earlier tutorial, a sample mean can have an approximately normal sampling distribution even when individual measurements are skewed, provided the sample is large enough for the Central Limit Theorem to apply.

Definition: Data are approximately normal when their distribution is reasonably close to a normal distribution in shape. In practice, we look for a roughly symmetric, single-peaked histogram without strong skewness or influential outliers, and we may also check whether a normal probability plot is approximately linear.

Reading a Histogram for Normality

A histogram is a first check because it shows the overall pattern of the observations. A normal-looking histogram has one main peak near its center, with the bars tapering off in a roughly balanced way toward both ends. The left and right sides need not be perfect mirror images; real samples vary. The question is whether the departures from symmetry are small enough that a normal model remains a reasonable approximation for the purpose at hand.

Look for features that may make a normal model questionable. A long tail on one side suggests skewness. Two or more distinct peaks suggest that the data may combine different groups or processes. A gap, a cluster separated from the rest, or an extreme value may also deserve attention. The sample size and the bin widths affect how a histogram looks, so a single unusual-looking bar is not automatically decisive. Consider the overall pattern and, when possible, compare it with another diagnostic.

What to look for in a histogram:
  • Center and peak: Is there one main cluster or peak?
  • Symmetry: Do the left and right sides have roughly similar shapes?
  • Tails: Do the tails taper in a reasonably balanced way, or is one much longer?
  • Unusual features: Are there gaps, multiple peaks, or possible outliers?

Symmetry alone is not enough. A flat, uniform distribution can be symmetric but is not bell-shaped. A histogram with two peaks may also look balanced from left to right without resembling one normal distribution. Use the full shape, not just whether the two sides appear similar.

Checking the Empirical Rule

The empirical rule, introduced in Using the Empirical Rule for Normal Areas, gives approximate benchmarks for normal distributions: about 68% of values are within one standard deviation of the mean, about 95% are within two, and about 99.7% are within three. For a data set, use the sample mean \(\bar{x}\) and sample standard deviation \(s\) to count observations in the corresponding intervals. Those percentages can indicate whether the data’s spread is consistent with a normal shape.

This is a check, not a pass-or-fail test. Sample percentages will not generally match 68%, 95%, and 99.7% exactly. A moderate difference can happen by chance, especially with a small sample. Also, the three percentages summarize only how many observations fall within certain distances of the mean; they do not describe the entire shape. A histogram or normal probability plot can reveal features that the percentages miss.

Formula: The three intervals to check are \(\bar{x}\pm s\), \(\bar{x}\pm2s\), and \(\bar{x}\pm3s\). For each interval, count the observations inside it and divide by the total sample size to find the observed proportion.

Be careful about what the histogram can establish. If a bin falls entirely inside an interval, all observations in that bin count as being inside. If a bin crosses an interval boundary, the histogram alone may not tell you how many observations from that bin fall inside. Use the raw data or a calculator’s one-variable statistics and list features when you need an exact count.

Using a Normal Probability Plot

A normal probability plot compares ordered data values with the values expected from a normal distribution. The plot can be made with the theoretical normal scores on one axis and the ordered observations on the other. If the data are approximately normal, the points tend to follow a roughly straight-line pattern. The line does not have to be perfect; small irregularities are expected in a sample.

Systematic departures from a line can signal that a normal model is questionable. A curve or pronounced bend may indicate skewness. Points that separate from the overall pattern at one or both ends may suggest unusually extreme observations or tails that differ from a normal distribution. A straight-looking middle section does not erase a strong departure in the tails: look across the whole plot.

Definition: A normal probability plot is a graph that compares ordered observed values with corresponding theoretical normal scores. A roughly straight pattern supports using a normal model as an approximation; a systematic curve or isolated points can count against it.

The plot is a visual diagnostic, not a proof. With a small sample, even data from a normal population may produce a noticeably uneven plot. With a large sample, small departures from normality may become visible. Interpret the plot alongside the histogram and the setting rather than using a rigid rule such as “every point must be on a line.”

A Practical Assessment

Use the checks together. Start with a histogram to see the overall shape and identify possible skewness, multiple peaks, or outliers. If the distribution looks reasonably symmetric and single-peaked, compare the observed percentages within one, two, and three standard deviations with the empirical-rule benchmarks. Then inspect a normal probability plot if one is available, paying attention to curvature and the tails.

1
Identify the data being assessed.
State what each observation measures and whether you are checking individual values or a statistic’s sampling distribution.
2
Inspect the histogram.
Describe its peak, symmetry, tails, and any gaps, clusters, or possible outliers.
3
Check additional evidence.
If appropriate, compare the observed within-standard-deviation percentages with the empirical rule or look for a roughly linear normal probability plot.
4
Give a qualified conclusion.
Explain whether the evidence makes a normal model reasonable, questionable, or unsuitable for the intended use. Do not claim that a check proves exact normality.

Worked Example: A Nearly Symmetric Set of Completion Times

Worked Example: A Nearly Symmetric Set of Completion Times

In a fictional fitness program, 100 participants record how many minutes they take to complete a particular course. The histogram is single-peaked and roughly symmetric. The sample mean is \(\bar{x}=12\) minutes and the sample standard deviation is \(s=2\) minutes. The counts in consecutive two-minute bins from 6 to 18 minutes are 3, 13, 34, 34, 13, and 3; no observations are below 6 or above 18 minutes. Assess whether a normal model is reasonable.

State. We are assessing the distribution of individual completion times, measured in minutes, for these 100 participants. We will consider the histogram and compare the observed within-standard-deviation percentages with the empirical rule.

Plan. A normal distribution should be single-peaked and approximately symmetric. The histogram description and balanced bin counts are relevant shape evidence. We will find the intervals \(\bar{x}\pm s\), \(\bar{x}\pm2s\), and \(\bar{x}\pm3s\), count observations inside each interval, and compare the resulting percentages with 68%, 95%, and 99.7%. Because the bin boundaries match the interval boundaries, these counts can be read from the stated bins.

Do. The three intervals are:

$$ \bar{x}\pm s=12\pm2=[10,14] $$
$$ \bar{x}\pm2s=12\pm4=[8,16] \qquad \bar{x}\pm3s=12\pm6=[6,18] $$

The bins from 10 to 14 minutes contain \(34+34=68\) participants. Thus, \(68/100=0.68\), or 68%, are within one standard deviation. The bins from 8 to 16 minutes contain \(13+34+34+13=94\) participants. Thus, \(94/100=0.94\), or 94%, are within two standard deviations. All 100 participants are between 6 and 18 minutes, so \(100/100=1.00\), or 100%, are within three standard deviations.

These observed percentages, 68%, 94%, and 100%, are close to the empirical-rule benchmarks of about 68%, 95%, and 99.7%. The bin counts are also balanced on either side of the center: 3 and 3, 13 and 13, and 34 and 34.

Conclude. The roughly symmetric, single-peaked histogram and the close empirical-rule percentages support treating these completion times as approximately normal. They do not prove that the population distribution is exactly normal; a normal probability plot would provide another check.

Worked Example: A Right-Skewed Distribution of Wait Times

Worked Example: A Right-Skewed Distribution of Wait Times

A fictional walk-in service records the waiting times, in minutes, for 50 visitors. A histogram has these counts: 25 visitors wait from 0 to less than 5 minutes, 15 wait from 5 to less than 10 minutes, 6 wait from 10 to less than 15 minutes, 3 wait from 15 to less than 20 minutes, and 1 waits from 20 to less than 25 minutes. Assess whether a normal model is a reasonable description of these data.

State. We are assessing the distribution of the 50 individual waiting times, in minutes.

Plan. We will describe the histogram’s peak, symmetry, and tails. We will also check that the listed frequencies sum to the stated sample size so the pattern is being interpreted consistently. A normal-looking histogram should be roughly symmetric, rather than having a much longer tail on one side.

Do. The counts sum to \(25+15+6+3+1=50\), as stated. The largest count is in the first bin, and the frequencies generally decrease as waiting times increase. The data extend through four bins above 5 minutes but only one bin below 5 minutes, since waiting times cannot be negative. This is a long tail to the right, not a balanced taper on both sides.

Conclude. The histogram is strongly right-skewed, so a normal model is not a reasonable description of these individual waiting times. A normal probability plot would be expected to show a systematic departure from a straight-line pattern, especially toward the upper tail. The fact that one can calculate a mean and standard deviation does not make the distribution normal.

Worked Example: Reading a Normal Probability Plot

Worked Example: Reading a Normal Probability Plot

A fictional sample contains nine measurements of a small component’s width, in millimeters. A normal probability plot places theoretical normal scores on the horizontal axis and the ordered measurements on the vertical axis. The plotted pairs, rounded, are shown below. Assess whether the plot supports a normal model.

Theoretical normal scoreOrdered width (mm)
-1.5934.2
-0.9674.8
-0.5895.2
-0.2825.6
0.0006.0
0.2826.4
0.5896.8
0.9677.2
1.5937.8

State. We are assessing the distribution of individual component widths, measured in millimeters, using the displayed normal probability plot.

Plan. A normal probability plot supports an approximate normal model when its points follow a roughly straight-line pattern. We will compare the observed widths with the theoretical scores across the center and both ends, looking for curvature or isolated points.

Do. As the theoretical score increases from \(-1.593\) to \(1.593\), the ordered widths increase steadily from 4.2 to 7.8 millimeters. The increases are fairly regular across the table: the points would follow a roughly straight upward pattern, with no single value separated sharply from the rest and no obvious bend in either tail.

Conclude. The normal probability plot is consistent with an approximately normal distribution of component widths. With only nine observations, the plot cannot establish normality; the small sample can show irregular patterns just by chance. The plot provides supportive evidence, not proof.

Common Mistakes and AP Exam Tips

  • Calling any symmetric histogram normal. Symmetry matters, but so do the single peak and overall bell-shaped pattern. A flat or two-peaked symmetric distribution is not normal-looking.
  • Treating the empirical rule as exact. The benchmarks are approximate. State the observed counts or percentages and describe whether they are reasonably close, rather than requiring exact equality.
  • Counting from bins that cross an interval boundary. A histogram may not reveal how many observations in a crossing bin are inside the interval. Use the raw data or calculate the count from the observations instead of guessing.
  • Letting a good middle hide a bad tail. A normal probability plot can look nearly linear in the center but curve at an end. Comment on the full pattern and any isolated points.
  • Claiming that normality has been proved. A histogram, empirical-rule comparison, or normal probability plot provides evidence about whether a normal approximation is reasonable. None proves that the population is exactly normal.
  • Confusing individual data with a sampling distribution. A histogram of individual measurements checks their distribution. The sampling distribution of \(\bar{x}\) may be approximately normal under the conditions discussed in Normal Models for Sample Means and Proportions, even if individual data are skewed.

A strong AP response names the feature that supports or challenges normality and connects that feature to the conclusion. For example: “The histogram is single-peaked and roughly symmetric, and the observed percentages within one, two, and three standard deviations are close to the empirical-rule benchmarks, so a normal model appears reasonable.” If the shape is skewed, identify the direction of skew and explain why that weakens the case for a normal model.

AP Exam Tip: Use cautious wording such as “the evidence is consistent with an approximately normal distribution” or “the strong right skew makes a normal model questionable.” Do not write that a visual check proves the data are normal.

Key Takeaway

A normal model is most credible when multiple features agree: the histogram is single-peaked and roughly symmetric, the empirical-rule percentages are reasonably close to their benchmarks, and the normal probability plot is roughly linear. A clear skew, multiple peaks, or strong curvature is evidence against using a normal model for the data being assessed.

Key takeaway: Check the whole shape, not a single statistic. Use histograms, empirical-rule percentages, and normal probability plots as complementary evidence, and make a qualified conclusion about whether a normal model is reasonable for the particular data or statistic.

Check Your Understanding

For each situation, describe what the evidence says about using a normal model.

  1. A histogram is single-peaked but has a much longer tail to the left. What feature makes a normal model questionable?
  2. A sample has 200 observations. There are 136 within one standard deviation of the mean, 188 within two, and 200 within three. Compare these percentages with the empirical-rule benchmarks.
  3. A normal probability plot is nearly straight in the middle but bends upward sharply at the largest observations. What should you mention in your assessment?
  4. Explain why a symmetric histogram with two distinct peaks is not necessarily approximately normal.
  5. What is one limitation of using the empirical rule to assess a distribution?