Check the Right Distribution
When checking conditions for mean inference, it is easy to look at the wrong thing. In “What to Do When Conditions Are Not Met,” you learned to respond to the specific condition that is unsupported or unclear. This tutorial focuses on a common source of mistakes: confusing the distribution of individual observations with the sampling distribution of the sample mean, or using the \(n\geq30\) rule when it does not apply.
A one-sample t procedure uses sample data to make an inference about a population mean \(\mu\). Its shape condition is not a demand that every sample look perfectly Normal. The concern is whether the sampling distribution of \(\bar{x}\) is reasonably modeled by a t procedure. The Normal/Large Sample condition provides evidence for that model: the population is approximately Normal, or the large-sample route applies. With a small sample, graphs of the individual observations help assess whether the population shape is plausibly suitable.
As explained in “Common Mistakes With Sampling Distributions of Means,” the sampling distribution of \(\bar{x}\) is not the same thing as the distribution of individual observations. In practice, a single sample does not let you directly inspect the sampling distribution. Instead, use the study information and the sample-size route, or examine a graph of the observed individual values when the sample is small.
What the n-at-Least-30 Rule Does—and Does Not—Say
“Using the n at Least 30 Rule Correctly” established the AP large-sample route: \(n\geq30\) supports using an approximately Normal model for the sampling distribution of \(\bar{x}\). It does not say that the individual observations become Normal, that every graph must look bell-shaped, or that any sample with at least 30 observations is automatically free of all concerns.
The rule must match the actual sample size. If \(n=18\), it is incorrect to cite \(n\geq30\), even if the sample mean and standard deviation seem reasonable. For \(n<30\), check whether the population is approximately Normal or use sample graphs to look for evidence about its shape. Strong skewness or a pronounced outlier in a small sample weakens the support for a t procedure.
For a larger sample, the large-sample route supports the sampling-distribution model, but it does not turn the observations into a Normal distribution. Extreme outliers or a very unusual shape still deserve attention, as discussed in “Robustness of t Procedures.” Also, the shape check is only one part of a condition audit. It does not replace checks of randomness or independence.
A Reliable Shape-Check Routine
For a small sample, the graphs and the context provide evidence, not proof, about the population shape. A dotplot, histogram, or boxplot can reveal skewness, gaps, clusters, and possible outliers. A Normal probability plot can help assess whether the sample pattern is roughly consistent with a Normal population. The earlier tutorials “Graphing Sample Data to Check Normality,” “Reading a Normal Probability Plot,” and “Identifying Skewness and Outliers in Small Samples” describe what to look for.
Do not judge shape by one feature alone. A sample may be somewhat asymmetric without showing strong skewness; a single extreme observation can matter more than mild irregularity. Explain what the display shows and connect that evidence to the sample size. If a value looks unusual, investigate whether it is an error, but do not delete a genuine observation just to make the graph look more Normal.
Check that the graph displays the individual measurements relevant to the mean procedure, and record the actual \(n\).
If \(n\geq30\), state that the large-sample route supports an approximately Normal sampling distribution. If \(n<30\), assess the population-shape evidence, using graphs of the sample when available.
Name features such as approximate symmetry, strong skewness, or a pronounced outlier. Do not claim to have inspected the sampling distribution of \(\bar{x}\) from one sample.
Check randomness and independence, including the 10% condition when sampling without replacement. Shape evidence cannot establish those design conditions.
Worked Examples
Worked Example: A Small Sample Does Not Meet the Large-Sample Route
A fictional community garden randomly selects 18 tomato plants from its 240 plants and records each plant’s fruit mass, in grams. A histogram shows a long right tail and one unusually large observation. The gardeners want a one-sample t interval for the mean fruit mass of all 240 plants.
State. Let \(\mu\) be the mean fruit mass, in grams, for the garden’s 240 tomato plants. We need to decide whether the conditions support using a one-sample t interval.
Plan. Check random selection, independence and the 10% condition, then assess the Normal/Large Sample condition. Since this is a small sample, use the graph of the individual fruit masses to look for strong skewness and outliers. Do not apply the large-sample route unless the actual sample size is at least 30.
Do. The plants were randomly selected, supporting the Random Condition for this garden. Since sampling was without replacement, check the 10% condition: \(18\leq0.10(240)=24\), so the sample is no more than 10% of the population. The large-sample route does not apply because \(18<30\). The histogram shows a long right tail and an unusually large observation, so the sample does not provide reassuring evidence that the population distribution is approximately Normal. The gardeners should check that the unusual value is recorded correctly, but should not remove it if it is a genuine measurement.
Conclude. The random-selection and 10% checks are supported, but the Normal/Large Sample condition is not supported by the evidence given: \(n=18\), and the sample is strongly right-skewed with a possible outlier. The gardeners should not justify the usual t interval by claiming \(n\geq30\). They could investigate the unusual measurement and seek more appropriate data, while reporting that this sample does not provide strong support for the proposed t inference.
Worked Example: The Graph Is Not a Graph of the Sample Mean
A fictional clinic randomly selects 16 patients and records the number of minutes each patient waited for an appointment. A student says, “The histogram of the sample means is right-skewed, so the Normal/Large Sample condition fails.” The student has only one set of 16 wait times and has made a histogram of those 16 individual values.
State. Let \(\mu\) be the mean appointment wait, in minutes, for the clinic’s patients. We need to evaluate whether the student’s claim describes the evidence correctly.
Plan. Clarify what the graph displays. A histogram of the 16 recorded wait times displays individual observations in this sample, not the sampling distribution of \(\bar{x}\). Because \(n<30\), use that sample graph as evidence about population shape, while checking the design conditions separately.
Do. The sample size is \(n=16<30\), so the large-sample route is not available. The histogram can be used to describe the observed individual wait times and to assess whether they provide support for an approximately Normal population distribution. But the student has not repeatedly drawn samples and calculated a mean for each one; therefore, the histogram is not a histogram of the sampling distribution of \(\bar{x}\). The description says the sample was randomly selected, supporting the Random Condition. The clinic’s patient population size and whether selection was without replacement are not given, so the 10% check cannot be completed from the information provided.
Conclude. The student should say that the graph shows the distribution of the 16 observed wait times, then describe its shape. The Normal/Large Sample condition is unclear unless the graph and other information provide evidence about population shape; it cannot be rejected simply by calling the graph a sampling distribution of \(\bar{x}\). The independence check is also incomplete without the relevant population and sampling details.
Worked Example: A Larger Sample Does Not Make the Data Normal
A fictional sports-science class randomly selects 36 runners from a large running club and records their weekly training distances. The histogram of the individual distances is moderately right-skewed, with no isolated extreme value. The class wants a one-sample t interval for the club’s mean weekly training distance.
State. Let \(\mu\) be the mean weekly training distance, in kilometers, for runners in the club. We need to assess whether the shape check supports a one-sample t interval.
Plan. Check the Random Condition, independence, the 10% condition if the sample was taken without replacement, and the Normal/Large Sample condition. Since the actual sample size is at least 30, use the large-sample route to assess the sampling distribution of \(\bar{x}\). Describe the observed shape without incorrectly claiming the individual distances are Normal.
Do. The runners were randomly selected, supporting the Random Condition. The club is described as large, but the exact population size and sampling details are not given, so we cannot show the 10% comparison from the information provided. The shape check has \(n=36\geq30\), so the large-sample route supports an approximately Normal model for the sampling distribution of \(\bar{x}\). The individual distances are moderately right-skewed; the large-sample route does not change that description. There is no isolated extreme value in the graph, which avoids an additional outlier concern.
Conclude. The sample-size route supports the Normal/Large Sample condition for the sampling distribution of \(\bar{x}\), even though the individual distances are right-skewed. The class should not write that the individual data are Normal. Before presenting the t interval as fully justified, it still needs to establish independence and, if sampling without replacement, check the 10% condition using the club’s population size.
Worked Example: Correcting an Incorrect Condition Statement
A fictional school surveys a random sample of 27 students about how many hours they spend on a project each week. The sample histogram is roughly symmetric, with no apparent outliers. A response says, “The data are Normal because \(n=27\geq30\), so a t interval is appropriate.”
State. Let \(\mu\) be the mean number of project hours per week for all students at the school. We need to revise the condition statement to match the sample size and available shape evidence.
Plan. Identify the incorrect numerical claim, then use the small-sample route. Describe the graph as evidence about population shape, rather than asserting that it proves the population is Normal. Also note that the shape check alone does not establish every condition.
Do. The numerical comparison is false: \(27<30\), so the large-sample route does not apply. The histogram is roughly symmetric and has no apparent outliers, which provides some support for an approximately Normal population shape, but a graph of one sample cannot prove the population is Normal. Random selection supports the Random Condition. To complete the independence check, the response would also need to establish that the observations are independent and check the 10% condition if the students were sampled without replacement from a finite school population.
Conclude. A more accurate statement is: “Because \(n=27<30\), the large-sample route is not available. The sample histogram is roughly symmetric with no apparent outliers, providing support for an approximately Normal population shape. Random selection supports the Random Condition; independence and the 10% condition should also be checked.” This statement uses the actual sample size and does not overstate what the graph shows.
Common Mistakes and AP Exam Tips
- Claiming \(n\geq30\) when it is not true. Compare the stated sample size with 30. If \(n=27\), write \(27<30\); then use population-shape evidence or sample graphs instead of the large-sample route.
- Calling individual data Normal because \(n\geq30\). The rule supports an approximately Normal sampling distribution of \(\bar{x}\), not a Normal distribution of the individual observations.
- Calling a sample histogram a sampling-distribution graph. A histogram of the observed values shows the individual data in one sample. State what is actually plotted before interpreting the shape.
- Assuming a graph proves the population shape. For a small sample, a graph provides evidence, not certainty. Use careful language such as “the sample appears roughly symmetric” rather than “the population is definitely Normal.”
- Using shape evidence to skip design checks. A reasonable-looking graph does not establish random selection, independence, or the 10% condition. Check each condition for its own reason.
- Declaring a small sample acceptable despite strong skewness or an outlier. With \(n<30\), strong skewness or a pronounced outlier weakens support for a t procedure. Do not cite robustness as a guarantee, or remove a genuine unusual value without justification.
A full-credit condition check names the relevant evidence and explains what it supports. For example: “Since \(n=18<30\), the large-sample route is not met. The sample histogram is strongly right-skewed with a pronounced high value, so it does not provide good support for an approximately Normal population shape.” This is more precise than “the data are not Normal,” and it avoids confusing individual values with the sampling distribution.
Check Your Understanding
For each situation, identify the shape-checking issue and state what a careful condition check should say.
- A sample has \(n=22\), but a student writes that the large-sample route applies because \(22\geq30\). What is wrong, and what should be checked instead?
- A histogram of 35 individual measurements is roughly bell-shaped. Does this graph display the sampling distribution of \(\bar{x}\)? Explain.
- A random sample of 14 observations is strongly left-skewed and has one pronounced low outlier. What does this suggest about using a t procedure?
- A sample has \(n=40\), and its individual observations are right-skewed. What does the large-sample route support, and what does it not say about the individual observations?
- A small sample’s graph looks roughly symmetric. Why should a response describe this as evidence rather than proof, and which other conditions still need checking?