Tutorials › AP Statistics › Central Limit Theorem Explained

Sampling distributions for means · Tutorial 607 of 1000

Central Limit Theorem Explained

See how sample means from a skewed population become approximately Normal as sample size grows, and practice using that model to find probabilities.

Intermediate 9 min read

What You'll Learn

  • State the Central Limit Theorem for a random sample from one population with finite mean and standard deviation
  • Distinguish an exactly Normal sampling distribution from an approximately Normal one
  • Find the approximate mean and standard deviation of a sample mean
  • Use the Central Limit Theorem to estimate probabilities for sample means from a skewed population
  • Check the sampling assumptions and explain why the needed sample size depends on population shape
  • Avoid treating the Central Limit Theorem as a guarantee that every sample mean is close to the population mean

Why Sample Means Can Be Normal Even When the Population Is Skewed

In “Sampling Distribution When the Population Is Normal,” we saw that a Normal population makes the sampling distribution of \(\bar{x}\) exactly Normal for any sample size, provided observations are independent. But many measurements are not Normally distributed. Wait times, for example, often have a long right tail: most waits may be moderate, while a few are unusually long.

The Central Limit Theorem (CLT) explains why the distribution of sample means can still be approximately Normal when samples are large enough. It concerns the sampling distribution of \(\bar{x}\), not the distribution of individual observations. A skewed population does not become Normal; rather, the means of repeated random samples from it tend to have a more nearly Normal distribution as the sample size increases.

Definition: If \(X_1,\ldots,X_n\) are a random sample from the same population with finite mean \(\mu\) and finite standard deviation \(\sigma\), then, as \(n\) increases, the sampling distribution of \(\bar{x}\) becomes approximately Normal. Its mean is \(\mu\), and its standard deviation is \(\sigma/\sqrt{n}\), provided the observations can be treated as independent.

The requirement that the observations come from the same population matters. It means they share a common distribution; in particular, they have the same population mean and standard deviation. Independence and matching means and standard deviations alone are not enough to state the usual CLT. The theorem also requires a population with finite mean and standard deviation.

For sampling without replacement from a finite population, observations are not literally independent. As covered in “The 10% Condition for Sample Means,” when the sample is less than 10% of the population, it is reasonable to treat the observations as approximately independent for the standard-deviation calculation. For an AP Statistics application, also describe how the sample was obtained and consider whether the population’s shape makes a Normal approximation reasonable.

$$ \text{Random sample from one population with finite }\mu\text{ and }\sigma \quad\Longrightarrow\quad \bar{X}\text{ is approximately }N\left(\mu,\frac{\sigma}{\sqrt{n}}\right) \text{ for sufficiently large }n $$

Here, the second parameter in \(N(\mu,\sigma/\sqrt{n})\) is the standard deviation. The center and spread come from earlier results: as in “Mean of the Sampling Distribution of x-bar,” the mean of \(\bar{x}\) is \(\mu\); as in “Standard Deviation of the Sample Mean,” its standard deviation is \(\sigma/\sqrt{n}\) when independence is appropriate. The CLT adds the important claim about the sampling distribution’s shape: for a sufficiently large sample, it is approximately Normal.

Key distinction: If the population is Normal and observations are independent, \(\bar{x}\) is exactly Normal for any positive \(n\). If the population is not Normal, the CLT may justify an approximately Normal sampling distribution for a sufficiently large \(n\). “Sufficiently large” depends in part on how strongly skewed or irregular the population is.

Using the CLT to Find Probabilities

Once the CLT makes a Normal approximation reasonable, calculate a probability about \(\bar{x}\) using a Normal model centered at \(\mu\) with standard deviation \(\sigma/\sqrt{n}\). Keep the sample size attached to the model: changing \(n\) changes the spread of sample means, even though their center remains \(\mu\).

1
Describe the population and sampling.
Identify the population mean and standard deviation, the sampling method, and whether independence is reasonable. Confirm that the data represent a random sample from the same population.
2
Assess the shape and sample size.
If the population is not Normal, decide whether the sample is large enough for the CLT approximation. Strong skewness or unusual features can require a larger sample; there is no single sample-size cutoff that guarantees success for every population.
3
Find the sampling distribution model.
Use mean \(\mu\) and standard deviation \(\sigma/\sqrt{n}\). State that the distribution of \(\bar{x}\) is approximately Normal, not exactly Normal.
4
Calculate and interpret.
Use normalcdf or a standardization calculation for the sample-mean values in the question. Conclude in context, with units, and describe the result as approximate.

The CLT does not say that every sample mean will be close to \(\mu\), or that a sample of a particular size must produce a Normal-looking distribution. It describes the shape of the distribution across repeated samples. The probability calculations account for the fact that sample means vary from sample to sample.

Worked Example: Mean Wait at a Service Desk

Suppose a hypothetical service desk has a right-skewed distribution of individual wait times, with population mean \(\mu=12\) minutes and population standard deviation \(\sigma=12\) minutes. Assume a random sample of \(n=64\) waits is selected, and the population is large enough that the 10% condition is met. Estimate the probability that the sample mean wait is more than 15 minutes.

State. Let \(\bar{x}\) be the mean wait time, in minutes, for the 64 sampled customers. We want \(P(\bar{x}>15)\).

Plan. The waits are sampled randomly from the same population, and the 10% condition supports treating the observations as approximately independent. The population has a finite mean and standard deviation, as specified. Although individual waits are right-skewed, a sample of 64 gives a basis for using the CLT approximation. This is an approximation, not a claim that the population or the sample of individual waits is Normal.

Do. The approximate sampling distribution of \(\bar{x}\) has mean 12 minutes and standard deviation:

$$ \sigma_{\bar{x}} = \frac{\sigma}{\sqrt{n}} = \frac{12}{\sqrt{64}} = \frac{12}{8} = 1.5\text{ minutes} $$

Standardize 15 minutes and find the upper-tail area:

$$ P(\bar{x}>15) \approx P\left(Z>\frac{15-12}{1.5}\right) = P(Z>2) \approx 0.0228 $$

A calculator check is \(\mathrm{normalcdf}(15,1\mathrm{E}99,12,1.5)\approx0.0228\), rounded to four decimal places.

Conclude. Under the stated random-sampling assumptions and CLT approximation, the probability that the mean wait for 64 sampled customers exceeds 15 minutes is about 0.0228, or 2.28%.

What Sample Size Does—and Does Not—Tell You

As \(n\) increases, \(\sigma/\sqrt{n}\) decreases, so sample means tend to be less variable. This spread formula was established in “How Sample Size Affects Variability of x-bar.” The CLT addresses a separate question: whether the distribution’s shape is approximately Normal. A larger sample generally helps, but the quality of the approximation also depends on the population’s shape.

A mildly skewed population may need a less substantial sample for its sample means to look approximately Normal than a population with extreme skewness or unusually influential values. The next tutorial considers how to judge whether a particular sample size is large enough. For now, do not treat any one number as a universal rule or use the CLT to claim exact Normality.

Also distinguish a probability about one observation from a probability about an average. For an individual wait, the population’s right-skewed distribution is relevant. For the mean of a sufficiently large random sample, the CLT supports using an approximately Normal model with a smaller standard deviation.

Worked Example: Mean Waits With a Larger Sample

Use the same hypothetical right-skewed wait-time population, with \(\mu=12\) minutes and \(\sigma=12\) minutes. A random sample of \(n=100\) waits is selected from a population large enough to satisfy the 10% condition. Estimate the probability that the sample mean is between 10 and 14 minutes.

State. Let \(\bar{x}\) be the mean wait time, in minutes, for the 100 sampled customers. We want \(P(10<\bar{x}<14)\).

Plan. The sample is random and comes from one population with finite mean and standard deviation. The 10% condition supports the independence approximation. Using the CLT, model the sampling distribution of \(\bar{x}\) as approximately Normal, centered at 12 minutes.

Do. First calculate the standard deviation of the sample mean:

$$ \sigma_{\bar{x}} = \frac{12}{\sqrt{100}} = 1.2\text{ minutes} $$

The standardized endpoints are:

$$ \frac{10-12}{1.2}\approx -1.6667, \qquad \frac{14-12}{1.2}\approx 1.6667 $$

Thus, using the Normal approximation:

$$ P(10<\bar{x}<14) \approx P(-1.6667<Z<1.6667) \approx 0.9044 $$

A calculator check is \(\mathrm{normalcdf}(10,14,12,1.2)\approx0.9044\), rounded to four decimal places.

Conclude. Under the stated sampling assumptions and CLT approximation, the probability that the mean wait for 100 sampled customers is between 10 and 14 minutes is about 0.9044, or 90.44%.

Reading the Model Carefully

In the examples, the Normal model describes \(\bar{x}\), not a customer’s individual wait. For instance, a right-skewed distribution of individual waits does not become symmetric just because 64 or 100 waits were collected. Instead, if we repeatedly took random samples of that size and plotted their means, the resulting sampling distribution would tend to be more nearly Normal.

The standard deviation \(\sigma/\sqrt{n}\) is the typical distance of sample means from \(\mu\) under the independence model. It is not a guarantee that a particular sample mean will be within that distance, and it is not the standard deviation of individual waits. Keep the parameter and statistic distinct: \(\mu\) and \(\sigma\) describe the population; \(\bar{x}\) and its sampling distribution describe sample means.

Worked Example: A Probability for a Sample of 36

For the same hypothetical population of right-skewed wait times, \(\mu=12\) minutes and \(\sigma=12\) minutes. Suppose a random sample of \(n=36\) waits is taken from a population large enough to satisfy the 10% condition. Using the CLT approximation, estimate the probability that the sample mean is greater than 14 minutes.

State. Let \(\bar{x}\) be the mean wait, in minutes, for the 36 sampled customers. The target is \(P(\bar{x}>14)\).

Plan. The observations form a random sample from the same population, with finite mean and standard deviation. The 10% condition supports approximate independence. We will use the CLT Normal approximation for this sample size, recognizing that its accuracy depends on how strongly skewed the population is.

Do. The standard deviation of the sample mean is:

$$ \sigma_{\bar{x}} = \frac{12}{\sqrt{36}} = \frac{12}{6} = 2\text{ minutes} $$

The value 14 is one standard deviation above the mean, so:

$$ P(\bar{x}>14) \approx P\left(Z>\frac{14-12}{2}\right) = P(Z>1) \approx 0.1587 $$

A calculator check is \(\mathrm{normalcdf}(14,1\mathrm{E}99,12,2)\approx0.1587\), rounded to four decimal places.

Conclude. Using the CLT approximation and the stated sampling assumptions, the probability that the mean wait for 36 sampled customers exceeds 14 minutes is about 0.1587, or 15.87%. Because the population is right-skewed, this approximation should be understood as an estimate rather than an exact probability.

Common Mistakes and AP Exam Tip

  • Applying the CLT to unrelated observations: The usual AP Statistics statement concerns a random sample from the same population. Do not claim that independence and equal means and standard deviations by themselves guarantee the CLT.
  • Saying the population becomes Normal: The CLT concerns the sampling distribution of \(\bar{x}\), not the distribution of individual measurements.
  • Calling the result exact: For a non-Normal population, the CLT gives an approximate Normal model. Exact Normality for every sample size comes from a Normal population, as in the previous tutorial.
  • Using \(\sigma\) instead of \(\sigma/\sqrt{n}\): The former is the standard deviation of individual observations; the latter is the standard deviation of sample means under the independence model.
  • Assuming one sample-size cutoff works for all populations: The amount of skewness and other population features affect how quickly the sampling distribution becomes approximately Normal. A sample-size claim should be justified in context.
  • Omitting the sampling assumptions: State that the sample is random and comes from the same population. For sampling without replacement, check and describe the 10% condition when appropriate.

For a full-credit explanation, identify the random-sampling and independence assumptions, name the CLT when the population is not Normal, give the approximate mean and standard deviation of \(\bar{x}\), show the probability calculation, and interpret the result in context. Use “approximately” when reporting a probability based on the CLT.

Key takeaway: For a random sample from one population with finite mean \(\mu\) and standard deviation \(\sigma\), the Central Limit Theorem says that the sampling distribution of \(\bar{x}\) is approximately Normal for a sufficiently large sample. Its mean is \(\mu\) and its standard deviation is \(\sigma/\sqrt{n}\), when independence is appropriate. The approximation is about sample means, and how large is “sufficiently large” depends partly on the population’s shape.

Check Your Understanding

Use the Central Limit Theorem and the sampling assumptions stated in each question.

  1. A population of individual delivery times is right-skewed, with mean 30 minutes and standard deviation 10 minutes. For a random sample of 25 deliveries, give the mean and standard deviation of the sampling distribution of \(\bar{x}\). Is its shape exactly Normal or approximately Normal under the CLT?
  2. Explain why “independent observations with the same mean and standard deviation” is not a complete statement of the usual CLT conditions. What population-based sampling description is appropriate?
  3. A random sample of 64 measurements comes from a population with mean 50 units and standard deviation 16 units. Assuming the sampling conditions are met, estimate \(P(\bar{x}>54)\) using the CLT.
  4. In a population of individual wait times, the distribution is strongly right-skewed. Explain why increasing the sample size helps the CLT approximation but does not make the individual wait-time distribution Normal.
  5. A student uses \(\sigma=12\) minutes as the standard deviation of \(\bar{x}\) for a sample of 36 waits. Identify the error and give the correct standard deviation of the sample mean.