Why Graph Small Samples?
In “Using the \(n\) at Least 30 Rule Correctly,” we saw that a sample with \(n<30\) does not meet the large-sample route for the Normal/Large Sample condition. In that case, graphs of the sample data can help us assess whether the population distribution might be approximately Normal. This tutorial focuses on reading that visual evidence carefully.
A graph displays the individual observations, not the sampling distribution of \(\bar{x}\). For a small sample, that distinction matters: a graph can show whether the sample has strong skewness, distinct clusters, or unusual values, but it cannot prove the population’s shape. A roughly symmetric sample with one main cluster and no clear outliers is generally more reassuring than a strongly skewed or irregular sample.
What to Look For in Each Graph
A dotplot is often especially useful for a small sample because it shows every observation. Each dot represents one value, and repeated values are stacked. Look for one main cluster, a roughly balanced shape on either side of the center, and any gaps or values far from the rest. A dotplot makes it easier to notice details that might be hidden when data are grouped into intervals.
A histogram groups values into bins, or intervals, and displays how many observations fall in each one. It can show the overall shape clearly, but the appearance depends partly on the chosen bin widths and boundaries. With a small sample, changing those choices can make the same data look more or less irregular. Avoid interpreting a histogram as if its bars show exact individual values.
A boxplot summarizes the distribution with the median, quartiles, and whiskers. It can help reveal asymmetry or a possible outlier, but it does not show every observation or reliably reveal multiple clusters. A boxplot that looks balanced does not rule out a gap or two separate groups in the data.
A roughly Normal shape is unimodal (has one main peak), approximately symmetric, and tapers toward both ends without clear outliers. For a small sample, the pattern need not look perfectly smooth. A few uneven stacks or bars can occur just because there are few observations. The important question is whether the overall pattern gives reasonable evidence for approximate Normality, not whether every dot forms a perfect bell shape.
A Practical Graph-Checking Process
Identify the quantitative variable, its units, and the population the sample is intended to represent.
For a small sample, use a dotplot when possible. A histogram or boxplot can add a useful view, but note what it may conceal.
State whether the pattern is roughly symmetric or skewed, whether it has one main cluster, and whether there are gaps or possible outliers.
For \(n<30\), explain whether the sample pattern is consistent with an approximately Normal population. Do not claim the graph proves the population’s shape.
This graph check addresses shape only. As covered in “Checking the Random Condition” and “Checking the 10% Condition for Independence,” randomness and independence need separate support. A favorable-looking graph cannot correct a biased selection method, and a random sample does not guarantee a Normal-looking sample.
Worked Examples
Worked Example: A Roughly Symmetric Sample of Batch Times
Suppose a quality technician randomly selects 17 batches from 500 batches and records the time, in minutes, each batch takes to pass through a cooling stage. The sample values are:
\(10,\ 12,\ 13,\ 14,\ 14,\ 15,\ 15,\ 16,\ 16,\ 16,\ 17,\ 17,\ 18,\ 18,\ 19,\ 20,\ 22\)
State. Let \(\mu\) be the mean cooling time for all 500 batches. We are checking whether the sample provides shape evidence for the Normal/Large Sample condition in a one-sample t procedure for \(\mu\).
Plan. Since \(n=17<30\), the large-sample route is not met. We will inspect the sample’s distribution for symmetry, a single main cluster, and possible outliers. We will also check that the sampling method and sample size support the separate random and independence conditions.
Do. A dotplot of these observations would have its highest stack at 16 minutes, with values generally becoming less frequent toward either end. The values are fairly balanced around 16: for example, 10 and 22 are equally far from 16, as are 14 and 18. There is one main cluster, and no observation stands far apart from the rest. This is a roughly symmetric, unimodal pattern without a clear outlier. It is consistent with an approximately Normal population, although 17 observations cannot establish that the entire population is Normal.
The technician selected batches at random, which supports the random condition. Because the sample was drawn without replacement, check the 10% condition:
The sample is 3.4% of the batches, less than 10%, so the 10% condition is met.
Conclude. The sample does not meet the \(n\geq30\) route, but its roughly symmetric, unimodal dotplot with no clear outliers gives reasonable evidence that the population distribution is approximately Normal. Together with the random selection and the 10% check, this supports the stated conditions for a one-sample t procedure, while not proving the population is Normal.
Worked Example: Right Skew and a Possible Outlier
A researcher randomly samples 16 repair requests from a large service center and records the time, in hours, until each request is completed. The times are:
\(4,\ 5,\ 5,\ 6,\ 6,\ 7,\ 7,\ 8,\ 8,\ 9,\ 10,\ 11,\ 12,\ 14,\ 18,\ 31\)
The researcher wants to assess whether the sample shape supports a t procedure for the population mean completion time.
A dotplot shows most observations between 4 and 14 hours, followed by a thinning set of larger values and then a gap before 31. This is a long right tail, not an approximately symmetric, bell-shaped pattern. A histogram with bins from 0 to less than 10, 10 to less than 20, 20 to less than 30, and 30 to less than 40 would have counts of 10, 5, 0, and 1, respectively. Those grouped counts also show a right-skewed pattern, though the dotplot makes the individual values clearer.
The boxplot offers a complementary check. Using the median-of-halves convention, the lower quartile is \(Q_1=6\), the upper quartile is \(Q_3=11.5\), and the interquartile range is \(11.5-6=5.5\). The upper outlier fence is:
Since 31 is above 19.75, the boxplot would flag it as a possible outlier. This rule identifies a value worth noticing; it does not prove that the value is an error or should be removed. The 18-hour observation is below the fence and is not flagged by this rule.
Conclusion. The sample has a pronounced right tail and a possible high outlier, and \(n=16<30\). These features do not provide reassuring evidence for an approximately Normal population. A t procedure’s Normal/Large Sample condition is not supported by this sample graph alone. Investigate whether the high value is valid and consider the data-collection context; do not delete an observation simply to make the graph look more Normal.
Worked Example: Why a Boxplot Is Not Enough
A transit analyst randomly selects 16 shuttle trips made during a particular period and records each trip’s loop time, in minutes. The times are:
\(10,\ 11,\ 11,\ 12,\ 12,\ 13,\ 13,\ 14,\ 22,\ 23,\ 23,\ 24,\ 24,\ 25,\ 25,\ 26\)
The trips were selected without replacement from 300 trips in that period. Assess the shape evidence for the population mean loop time.
State. Let \(\mu\) be the mean loop time for all trips in the specified period. The sample size is \(n=16\), so we need shape evidence rather than the large-sample route.
Plan. We will compare a dotplot and a boxplot, since the two displays emphasize different features. We will also check the random and 10% conditions separately.
Do. The dotplot has one cluster from 10 to 14 minutes and another from 22 to 26 minutes, with no observations from 15 through 21. The gap and two distinct clusters are not consistent with a single, approximately Normal mound. The pattern could reflect two types of trips or operating conditions; the context should be checked before treating the trips as one homogeneous group.
A boxplot summarizes the middle and overall spread, but it does not display the gap between these two clusters as clearly. Here, a boxplot alone could make the distribution seem like one broad spread of values. This is why a dotplot is valuable for small samples: it preserves the individual observations and can reveal structure hidden by a summary graph.
The trips were randomly selected, supporting the random condition. The sample fraction is:
Since \(5.33\%\) is below 10%, the 10% condition is met.
Conclude. Although the random and 10% conditions are supported, the sample’s two clusters and gap do not provide evidence for an approximately Normal population. The shape condition is not supported by this graph. The analyst should also investigate whether the trips combine distinct operating conditions before proceeding with inference about one population mean.
Common Mistakes and AP Exam Tips
- Claiming a sample graph proves population Normality. A graph shows the observed sample, not every member of the population. Say that the pattern is “consistent with” or “provides evidence for” approximate Normality; do not say it proves the population is Normal.
- Calling any uneven pattern non-Normal. Small samples naturally produce irregular stacks and bars. Focus on substantial features such as strong skewness, multiple clusters, or a clear outlier—not minor imperfections.
- Relying only on a boxplot. A boxplot is useful for spread and possible outliers, but it can hide clusters and gaps. For a small sample, check a dotplot or another display that shows individual values when possible.
- Ignoring histogram bin choices. A histogram’s appearance can change with its bin widths and boundaries. If the shape seems important, compare with a dotplot rather than treating one bin arrangement as decisive.
- Removing a flagged point automatically. A boxplot flag is a prompt to investigate, not proof of a recording error. Keep valid observations unless there is a defensible reason to exclude them.
- Mixing up individual data and \(\bar{x}\). These graphs describe individual observations. For \(n<30\), they help assess whether the population may be approximately Normal, which in turn supports a model for the sampling distribution of \(\bar{x}\).
- Letting the shape check replace other conditions. State how the sample was collected and address independence separately. A favorable shape does not establish randomness or eliminate bias.
A strong AP response names the sample size, describes specific graph features, and links those features to the condition with appropriately cautious wording. For example: “Because \(n=16<30\), the large-sample route is not met. The dotplot is roughly symmetric, has one main cluster, and shows no clear outliers, so it provides evidence that the population is approximately Normal; it does not prove this.” If the graph instead shows strong skewness or separated clusters, state that the sample does not provide reassuring evidence for the condition.
Check Your Understanding
Use the graph features described in each question to assess the shape evidence. Explain what you can and cannot conclude.
- A sample of 18 randomly selected garden plots has a dotplot with one roughly symmetric cluster and no clear outliers. Does the sample meet the \(n\geq30\) route? What evidence does the graph provide?
- A histogram of 15 observations looks strongly left-skewed, but a boxplot shows no flagged points. Which shape evidence should you report, and does the absence of a flagged point establish approximate Normality?
- Why might a dotplot reveal useful information that a boxplot of the same small sample does not?
- A boxplot flags one high observation using the \(1.5(\text{IQR})\) rule. What does the flag mean, and what should you investigate before deciding what to do with the value?
- A sample graph looks approximately Normal, but the observations came from volunteers rather than a random sample. Which condition does the graph address, and which concern remains?