Tutorials › AP Statistics › Common Mistakes With Sampling Distributions of Means

Sampling distributions for means · Tutorial 618 of 1000

Common Mistakes With Sampling Distributions of Means

Practice identifying the correct distribution and spread for a sample mean, and decide when a Normal approximation is justified.

Intermediate 9 min read

What You'll Learn

  • Distinguish the distribution of individual values from the sampling distribution of the sample mean.
  • Choose between the population standard deviation and the standard deviation of the sample mean.
  • Explain what the Central Limit Theorem does—and does not—say about sample means.
  • Check independence and assess whether a Normal model for sample means is appropriate.
  • Identify common errors and write accurate interpretations in context.

Three Different Questions Can Look Surprisingly Similar

In “Does a Larger Sample Reduce Bias in the Mean,” we separated variability from bias. Here, the focus is on reading and using the sampling distribution of \(\bar{x}\) correctly. Many errors start when a problem mentions a population, a sample, and a mean without making clear which one is being described.

Ask first: Is the question about one individual value, represented by the random variable \(X\), or about a sample mean, represented by \(\bar{x}\)? Then identify the relevant spread. Individual values vary with population standard deviation \(\sigma\). For independent observations, sample means vary with standard deviation \(\sigma/\sqrt{n}\). These spreads have the same units, but they describe different distributions.

A third question is whether the distribution of \(\bar{x}\) is Normal or approximately Normal. The population itself does not have to be Normal for a Normal approximation to the sampling distribution to be useful. But the Central Limit Theorem (CLT) is not a rule that makes every small-sample distribution Normal.

Quick check: Name the random quantity before choosing a distribution. For one observation, use the population distribution and spread \(\sigma\). For a sample mean, use the sampling distribution, centered at \(\mu\), with spread \(\sigma/\sqrt{n}\) when the independence conditions support that formula.

Mistake 1: Using \(\sigma\) for a Sample Mean

The population standard deviation \(\sigma\) describes how far individual observations tend to be from the population mean \(\mu\). It is not, by itself, the standard deviation of the sample mean. As established in “Standard Deviation of the Sample Mean,” independent observations from a population with standard deviation \(\sigma\) give

$$ \sigma_{\bar{x}}=\frac{\sigma}{\sqrt{n}}. $$

This quantity, \(\sigma_{\bar{x}}\), is the standard deviation of the sampling distribution of \(\bar{x}\). It is also called the standard error when referring to the typical variability of sample means. Increasing \(n\) makes this spread smaller; using \(\sigma\) instead of \(\sigma/\sqrt{n}\) ignores the fact that averaging observations reduces variability.

A useful unit check is not enough to catch this mistake: both \(\sigma\) and \(\sigma_{\bar{x}}\) have the original variable’s units. Instead, write the formula and identify the random quantity. If the question concerns individual observations, the spread is \(\sigma\). If it concerns averages of samples of size \(n\), the spread is \(\sigma/\sqrt{n}\).

Conditions: Use \(\sigma_{\bar{x}}=\sigma/\sqrt{n}\) when observations are independent. For a simple random sample without replacement from a finite population, check the 10% condition, as described in “The 10% Condition for Sample Means.” If the sample is not random or independence is not reasonable, do not assume this standard deviation formula applies without qualification.

Mistake 2: Confusing the Population With the Sampling Distribution

The population distribution describes individual values \(X\). The sampling distribution of \(\bar{x}\) describes the possible sample means from all samples of the same size, selected in the same way. A sample mean is not another individual observation, and a statement about the shape of the population is not automatically a statement about the shape of the sampling distribution.

The centers and spreads also refer to different things. Under the sampling conditions covered in “Mean of the Sampling Distribution of x-bar” and “Standard Deviation of the Sample Mean,” the sampling distribution of \(\bar{x}\) has mean \(\mu\) and standard deviation \(\sigma/\sqrt{n}\). The population distribution of \(X\), by contrast, has mean \(\mu\) and standard deviation \(\sigma\). The matching centers do not make the distributions identical.

This distinction matters when finding probabilities. \(P(X>c)\) is the probability that one observation exceeds \(c\). \(P(\bar{x}>c)\) is the probability that a sample mean exceeds \(c\). As explained in “Comparing Individual Values and Sample Means,” these are different questions and generally have different answers.

Mistake 3: Treating the CLT as a Guarantee for Any Sample Size

The CLT concerns the sampling distribution of the sample mean. For independent observations from the same population with finite mean \(\mu\) and finite standard deviation \(\sigma\), it says that as the sample size increases, the distribution of \(\bar{x}\) becomes approximately Normal, with mean \(\mu\) and standard deviation \(\sigma/\sqrt{n}\). The population itself does not have to become Normal.

The phrase “as the sample size increases” matters. The CLT does not say that a sample size of 30 always produces an adequate approximation, nor does it make the sampling distribution exactly Normal at every \(n\). How large a sample needs to be depends on the population’s shape. A population that is roughly symmetric and without outliers may need a smaller sample than one that is strongly skewed or has extreme values. “Is n = 30 Large Enough for the CLT” develops this judgment further.

Also check whether the observations are random and independent, and whether they come from the same population. For sampling without replacement from a finite population, check the 10% condition. Those checks address whether the sampling model is appropriate; the CLT addresses the shape of the sampling distribution as the sample size grows.

Key distinction: A Normal population makes the sampling distribution of \(\bar{x}\) Normal for any sample size when observations are independent. With a non-Normal population, a Normal model for \(\bar{x}\) may be an approximation supported by a sufficiently large sample and suitable conditions; it is not a claim that individual observations are Normal.

Worked Examples: Diagnose the Distribution Before Calculating

Worked Example: Individual Scores Versus the Mean Score

Suppose individual scores from an ongoing assessment process follow a Normal population with mean \(\mu=72\) points and standard deviation \(\sigma=15\) points. Independent random samples of \(n=25\) scores are collected. Compare the probability that one score exceeds 75 points with the probability that a sample mean exceeds 75 points.

State. The first probability concerns the individual random variable \(X\). The second concerns the sample mean \(\bar{x}\).

Plan. For one score, use the population standard deviation of 15 points. For the mean of 25 independent scores, use \(\sigma_{\bar{x}}=\sigma/\sqrt{n}\). Because the individual population is stated to be Normal and observations are independent, the sampling distribution of \(\bar{x}\) is also Normal. No 10% condition is needed for this independent production-process model.

Do. For one score, standardize 75 using the population standard deviation:

$$ z=\frac{75-72}{15}=0.20. $$

Thus, \(P(X>75)=P(Z>0.20)\approx0.4207\), rounded to four decimal places.

For a sample mean, first calculate its standard deviation:

$$ \sigma_{\bar{x}}=\frac{15}{\sqrt{25}}=\frac{15}{5}=3\text{ points}. $$

Then standardize 75 using the standard deviation of the sample mean:

$$ z=\frac{75-72}{3}=1.00. $$

Therefore, \(P(\bar{x}>75)=P(Z>1.00)\approx0.1587\), rounded to four decimal places. As a check, the two z-scores use the same difference, \(75-72=3\), but divide by different spreads: \(3/15=0.20\) for one score and \(3/3=1.00\) for the sample mean.

Conclude. The probability that one score exceeds 75 is about 0.4207, while the probability that the mean of 25 scores exceeds 75 is about 0.1587. The sample mean is less variable than an individual score, so 75 is farther above the center when measured in standard deviations of the sampling distribution.

Worked Example: Using the CLT With a Right-Skewed Population

A hypothetical city tracks the number of minutes residents spend waiting for a local service. The population is right-skewed, with mean \(\mu=30\) minutes and standard deviation \(\sigma=20\) minutes. A simple random sample of 100 residents is drawn from a population of 5,000. Estimate the probability that the sample mean wait exceeds 34 minutes.

State. The random quantity is the sample mean wait \(\bar{x}\), not one resident’s wait \(X\).

Plan. The sample is random. Check the 10% condition: \(100<0.10(5000)=500\), so the condition is satisfied and treating observations as independent is reasonable. The population is described only as right-skewed, and \(n=100\); without more information about the degree of skewness or outliers, the adequacy of a Normal approximation cannot be assessed from the stated information alone. For illustration, we will proceed using a Normal approximation, but this is an assumption rather than a conclusion supported by the description. We do not assume that individual wait times are Normal.

Do. The estimated center and standard deviation of the sampling distribution are

$$ \mu_{\bar{x}}=\mu=30\text{ minutes} \qquad\text{and}\qquad \sigma_{\bar{x}}=\frac{20}{\sqrt{100}}=2\text{ minutes}. $$

Standardize the cutoff of 34 minutes:

$$ z=\frac{34-30}{2}=2.00. $$

Using a Normal model, \(P(\bar{x}>34)\approx P(Z>2.00)=0.0228\), rounded to four decimal places. A second check is that 34 is exactly two standard errors above the sampling-distribution mean: \(30+2(2)=34\).

Conclude. Under the stated sampling conditions and using the CLT approximation, the probability that the mean wait for a random sample of 100 residents exceeds 34 minutes is about 0.0228. This calculation does not say that the probability one resident waits more than 34 minutes is 0.0228; that would be a question about \(X\), not \(\bar{x}\).

Worked Example: Why a Small Sample May Not Be Normal

Consider a population in which a randomly selected measurement \(X\) is 0 with probability \(3/4\) and 12 with probability \(1/4\). Independent observations are taken with replacement. A student claims that the sample mean for \(n=4\) is approximately Normal because of the CLT. Assess the claim and find the center and standard deviation of the sample mean.

State. The population has two possible values and is strongly right-skewed. The question concerns the sampling distribution of \(\bar{x}\) for a small sample, \(n=4\).

Plan. First find the population mean and standard deviation, then use the formula for the standard deviation of the sample mean. Independence is built into the with-replacement sampling. To assess the Normal claim, consider both the population’s shape and whether \(n=4\) is large enough for the CLT approximation. The CLT describes what happens as \(n\) increases; it does not guarantee a Normal approximation at this sample size.

Do. The population mean is

$$ \mu=\frac{3}{4}(0)+\frac{1}{4}(12)=3. $$

The population variance and standard deviation are

$$ \sigma^2=\frac{3}{4}(0-3)^2+\frac{1}{4}(12-3)^2 =\frac{3}{4}(9)+\frac{1}{4}(81)=27, \qquad \sigma=\sqrt{27}\approx5.196. $$

For \(n=4\), the sampling distribution is centered at \(\mu=3\), and its standard deviation is

$$ \sigma_{\bar{x}}=\frac{\sqrt{27}}{\sqrt{4}} =\frac{\sqrt{27}}{2} \approx2.598. $$

The same result follows by squaring the standard deviation: \(\sigma_{\bar{x}}^2=27/4=6.75\), and \(\sqrt{6.75}\approx2.598\). The sample mean can take only the values 0, 3, 6, 9, or 12, depending on how many of the four observations equal 12. That discrete sampling distribution is not itself a smooth Normal distribution.

Conclude. The student’s claim is not justified. The sample mean has center 3 and standard deviation about 2.598, but \(n=4\) is too small to rely on the CLT approximation for this strongly right-skewed population. The calculation of the center and standard deviation remains useful; it does not, by itself, establish a Normal shape.

Common Mistakes and AP Exam Tip

  • Using \(\sigma\) for \(\bar{x}\). This treats a sample mean like one observation. Write \(\sigma_{\bar{x}}=\sigma/\sqrt{n}\) and substitute the sample size when the question concerns means.
  • Using \(\sigma/\sqrt{n}\) for an individual \(X\). The sample-size adjustment belongs to the sampling distribution of the mean, not to the population distribution of individual values.
  • Calling the population Normal because of the CLT. The CLT concerns the distribution of sample means, not the distribution of individual observations.
  • Claiming the CLT guarantees Normality for \(n=30\). The required sample size depends on the population shape. Strong skewness or outliers can require a larger sample for a good approximation.
  • Checking the shape but ignoring the design. A large sample does not fix a nonrandom or dependent sampling process. Check random selection and independence as well as the conditions for a Normal approximation.
  • Giving a probability without identifying the variable. State whether the probability is about \(X\) or \(\bar{x}\), and interpret it in context. A complete answer makes clear whether it describes individuals or sample means.

For full-credit communication, name the random quantity, identify the distribution being used, show the relevant standard deviation, and explain why a Normal model is exact or approximately appropriate. If the sample is drawn without replacement from a finite population, explicitly check the 10% condition. If using the CLT, refer to the sampling distribution of the mean and avoid claiming that the population itself is Normal.

Key takeaway: Keep the individual distribution, the sampling distribution of \(\bar{x}\), and the conditions for a Normal model separate. Use \(\sigma\) for individual values and \(\sigma/\sqrt{n}\) for sample means under appropriate independence conditions. The CLT supports an approximation as sample size increases; it is not a guarantee for every population and sample size.

Check Your Understanding

For each question, identify the random quantity before choosing its distribution or spread.

  1. A population has standard deviation 18 units. What is the standard deviation of the sample mean for independent samples of size 36?
  2. A student uses \(\sigma/\sqrt{n}\) to find the probability that one randomly selected individual exceeds a cutoff. What is the error?
  3. A population is strongly right-skewed and a sample has size 5. Explain why the CLT alone does not justify treating the sampling distribution of the mean as Normal.
  4. A simple random sample of 80 is drawn without replacement from a population of 600. Check the 10% condition and state what it suggests about independence.
  5. Explain the difference between saying “individual measurements are Normal” and saying “the sampling distribution of the sample mean is approximately Normal.”