Tutorials › AP Statistics › Sampling Distribution of a Mean From a Skewed Population

Sampling distributions for means · Tutorial 614 of 1000

Sampling Distribution of a Mean From a Skewed Population

See how repeated sample means become less skewed and less variable as sample size increases, while their center remains at the population mean.

Intermediate 9 min read

What You'll Learn

  • Describe the sampling distribution of a mean when the population is strongly right-skewed.
  • Simulate repeated sample means for sample sizes 5, 15, and 40.
  • Explain how the sampling distribution’s shape changes as sample size increases.
  • Calculate and compare the center and standard deviation of the sample means.
  • Explain why a larger sample does not automatically guarantee an exactly Normal sampling distribution.

When the Population Is Strongly Right-Skewed

In “Central Limit Theorem Explained,” you learned that the sampling distribution of \(\bar{x}\) tends to become more nearly Normal as the sample size increases, provided observations are a random sample from the same population with a finite mean and standard deviation. How does this change look when the population is strongly skewed to the right? We can investigate by describing or simulating sample means at several sample sizes.

In this tutorial, we will use an invented, exponential-like population: most individual values are small, while a few are much larger. This population is discrete rather than a continuous exponential model, but it has the long right tail that makes it useful for studying the same shape changes. We will compare sampling distributions for \(n=5\), \(n=15\), and \(n=40\).

Definition: The sampling distribution of \(\bar{x}\) is the probability distribution of the sample means from all possible samples of a fixed size, selected using the same sampling method from the same population. Its shape describes how those means are distributed, not how individual population values are distributed.

That distinction matters. When the population is right-skewed, a small sample can include a very large observation, pulling its mean to the right. As the sample size increases, each individual observation has less influence on the mean, and the sampling distribution typically becomes less skewed. Its center and spread can be described separately from its shape.

An Exponential-Like Population to Study

Suppose each observation is selected independently from a large population with the following distribution. Imagine a bag of 100 equally likely tickets: 40 show 0, 30 show 10, 20 show 20, and 10 show 40. The values are measurements in minutes, and the ticket counts represent the population proportions.

Population value \(X\) (minutes)Population proportion
00.40
100.30
200.20
400.10

The probabilities decline as the value increases, with a few observations far above the most common values. That creates a pronounced right tail. The population mean is

$$ \mu=0(0.40)+10(0.30)+20(0.20)+40(0.10)=11\text{ minutes}. $$

For the population standard deviation, first calculate the variance from the weighted squared deviations:

$$ \sigma^2=(0-11)^2(0.40)+(10-11)^2(0.30)+(20-11)^2(0.20)+(40-11)^2(0.10)=149\text{ minutes}^2, $$ $$ \sigma=\sqrt{149}\approx12.2066\text{ minutes}. $$

The individual-value distribution is not symmetric: its right tail reaches 40 minutes, well above its mean of 11 minutes. Because the population has a finite mean and standard deviation, the center and standard deviation of sample means can be calculated when the sample observations are independent and come from this same population.

Formula: For independent observations from the same population with finite mean \(\mu\) and standard deviation \(\sigma\), the sampling distribution of \(\bar{x}\) has center \(\mu_{\bar{x}}=\mu\) and standard deviation \(\sigma_{\bar{x}}=\sigma/\sqrt{n}\). These formulas depend on the observations sharing that population’s mean and standard deviation.

How to Simulate the Sampling Distributions

A simulation makes the meaning of a sampling distribution concrete. For each chosen sample size, repeatedly draw a sample, calculate its mean, and record that mean. The collection of recorded means approximates the sampling distribution for that sample size.

1
Set up the population model.
Use the ticket proportions in the table. For each draw, select one of the 100 tickets at random and record its value.
2
Choose a sample size.
Begin with \(n=5\). Make five independent draws with replacement, so the same ticket value can occur more than once.
3
Calculate and record \(\bar{x}\).
Add the five values and divide by 5. Repeat the sampling process many times, recording one mean for every sample.
4
Repeat for \(n=15\) and \(n=40\).
Keep the population model and number of repetitions the same. Change only the sample size, then compare the three histograms of sample means.

For a physical simulation, numbered tickets or a random-number generator could select a ticket from 1 to 100, with the ranges assigned to the four values. Drawing with replacement supports independence. If instead we sampled without replacement from a finite population, we would need to consider the 10% condition, as discussed in “The 10% Condition for Sample Means.”

A real simulation produces slightly different histograms each time because of random variation. The important comparison is the overall pattern, not the exact count in a particular bin. The theoretical center and standard deviation give a useful check on what the simulated distributions should show.

Worked Examples: Following Shape, Center, and Spread

Worked Example: One Simulated Sample Mean

A simulated sample of five values from the population is \(0, 10, 10, 20,\) and \(40\) minutes. Find its mean and explain what this one result does—and does not—tell us.

State. The statistic is \(\bar{x}\), the mean of these five observations. It is one possible value from the sampling distribution for \(n=5\).

Plan. Add the five observed values and divide by the sample size. This is a sample drawn from the stated population model; the calculation describes this sample only, not all possible samples.

Do. The sum is \(0+10+10+20+40=80\) minutes, so

$$ \bar{x}=\frac{80}{5}=16\text{ minutes}. $$

As a check, the five values can be grouped as \(0+20+20+40=80\), which gives the same total and mean of 16 minutes.

Conclude. This sample mean is 16 minutes, which is above the population mean of 11 minutes. The large observation of 40 minutes pulls this particular mean upward. One sample mean does not show the whole sampling distribution; repeated samples are needed to see its overall shape and variability.

Worked Example: Compare the Standard Deviations for \(n=5\), \(15\), and \(40\)

For the exponential-like population, find the standard deviation of the sampling distribution of \(\bar{x}\) for each of the three sample sizes. Explain what the comparison predicts.

State. The population has \(\mu=11\) minutes and \(\sigma=\sqrt{149}\approx12.2066\) minutes. We are comparing the sampling distributions of the mean for \(n=5\), \(n=15\), and \(n=40\).

Plan. Each observation is an independent draw from the same population because the simulation uses replacement. The population has finite \(\mu\) and \(\sigma\). Therefore, use \(\mu_{\bar{x}}=\mu\) and \(\sigma_{\bar{x}}=\sigma/\sqrt{n}\) for each sample size.

Do. The center is 11 minutes for all three distributions. Their standard deviations are:

$$ n=5:\quad \sigma_{\bar{x}}=\frac{\sqrt{149}}{\sqrt{5}}=\sqrt{\frac{149}{5}}\approx5.4589\text{ minutes}, $$ $$ n=15:\quad \sigma_{\bar{x}}=\frac{\sqrt{149}}{\sqrt{15}}=\sqrt{\frac{149}{15}}\approx3.1517\text{ minutes}, $$ $$ n=40:\quad \sigma_{\bar{x}}=\frac{\sqrt{149}}{\sqrt{40}}=\sqrt{\frac{149}{40}}\approx1.9300\text{ minutes}. $$

As a check, calculate each standard deviation by first dividing the variance by \(n\), then taking its square root: \(\sqrt{149/5}\approx5.4589\), \(\sqrt{149/15}\approx3.1517\), and \(\sqrt{149/40}\approx1.9300\). The results match.

Conclude. All three sampling distributions are centered at the population mean of 11 minutes, while their spreads decrease as \(n\) increases. Sample means from samples of 40 tend to vary less around 11 minutes than means from samples of 5. These standard deviations describe the spread of sample means, not of individual observations.

Worked Example: Describe the Shape Changes

Suppose histograms are made from 10,000 simulated sample means at each of \(n=5\), \(n=15\), and \(n=40\), using the ticket population. Describe the expected shape changes without treating one random simulation’s exact histogram as a guaranteed result.

State. We are comparing the shapes of three simulated sampling distributions of \(\bar{x}\). The population is strongly right-skewed, and each sample consists of independent draws from the same population.

Plan. Use the population shape and the Central Limit Theorem to describe the expected progression. The theorem predicts that the sampling distribution becomes more nearly Normal as \(n\) grows; it does not say that it is exactly Normal for every finite sample size.

Do. For \(n=5\), a mean can be pulled noticeably to the right by one or more large observations, so the histogram of means is expected to remain right-skewed. For \(n=15\), the influence of any one observation is smaller, and the distribution of means is expected to be less skewed. For \(n=40\), the sampling distribution is expected to look more mound-shaped and roughly symmetric than it does at \(n=5\). Its center remains 11 minutes, and its standard deviation is about 1.9300 minutes, compared with 5.4589 minutes at \(n=5\).

The histogram need not become perfectly symmetric. Since the population values are discrete, sample means also take discrete values, and a particular simulated histogram can have irregularities. The mean is also bounded between 0 and 40 minutes, so the sampling distribution is not exactly an unrestricted Normal distribution.

Conclude. The expected pattern is a right-skewed sampling distribution at \(n=5\), less skew at \(n=15\), and a more nearly symmetric distribution at \(n=40\), with all three centered at 11 minutes and progressively smaller spread. These are predictions about the overall pattern; random simulation results will vary.

What Changes—and What Stays the Same

For independent samples from the same population, the mean of the sampling distribution stays at \(\mu\), regardless of sample size. In the example, all sample sizes have sampling-distribution center 11 minutes. A larger sample does not systematically make \(\bar{x}\) larger or smaller; it makes the sample mean less variable from sample to sample.

The standard deviation \(\sigma/\sqrt{n}\) also gets smaller as \(n\) increases. Here, moving from \(n=5\) to \(n=40\) multiplies the sample size by 8, so the standard deviation of \(\bar{x}\) is divided by \(\sqrt{8}\). Numerically, \(5.4589/\sqrt{8}\approx1.9300\) minutes, consistent with the direct calculation for \(n=40\).

Shape changes are not captured by the center and standard deviation alone. A distribution can have a known center and spread while still being strongly skewed. The CLT describes a tendency toward a Normal shape as sample size grows, but how quickly that happens depends on the population. A strongly skewed population may need a larger sample than a roughly symmetric population before its sample means look approximately Normal.

Therefore, do not use a fixed rule such as “\(n=30\) is always enough.” As discussed in “Is \(n=30\) Large Enough for the CLT,” judge the population’s skewness and look for the sampling distribution to become reasonably symmetric. In this example, \(n=40\) should be more nearly Normal than \(n=5\), but that is not a guarantee of exact Normality.

Common Mistakes and AP Exam Tip

  • Confusing the population with the sampling distribution: The individual-value population is strongly right-skewed. The sampling distribution is made of sample means and becomes less skewed as \(n\) increases. Describe which distribution you mean.
  • Claiming the center moves as \(n\) grows: For independent observations from the same population, the center remains \(\mu\). Increasing \(n\) reduces the spread; it does not change the population mean.
  • Saying the CLT makes every sample mean Normal: The result is that the sampling distribution becomes more nearly Normal as \(n\) increases, not that it is exactly Normal at any chosen sample size.
  • Using \(\sigma\) instead of \(\sigma/\sqrt{n}\) for sample means: \(\sigma\) describes individual observations. The sampling-distribution standard deviation is \(\sigma/\sqrt{n}\) under the stated independence and common-population assumptions.
  • Treating one simulated histogram as definitive: Random simulations vary. Compare the overall shape across many repetitions and connect the results to the expected center and spread.
  • Forgetting the sampling assumptions: State that observations are independent and come from the same population. For a finite population sampled without replacement, check the 10% condition when applying the independence-based formula.

For full-credit communication, name the statistic \(\bar{x}\), identify the population shape, and compare the sampling distributions by shape, center, and spread. A strong description says that the distribution of means becomes less skewed and narrower as sample size increases, while remaining centered at the population mean; it avoids claiming that the distribution is exactly Normal.

Key takeaway: With independent observations from the same right-skewed population, the sampling distribution of \(\bar{x}\) stays centered at \(\mu\) and has standard deviation \(\sigma/\sqrt{n}\). As \(n\) increases, it becomes less skewed and more nearly Normal, but the rate of change depends on the population.

Check Your Understanding

Use the exponential-like population and ideas in this tutorial to answer the questions.

  1. For independent samples of size 40 from the ticket population, what is the center of the sampling distribution of \(\bar{x}\)? Explain why it does not change from the center for samples of size 5.
  2. Calculate the standard deviation of the sampling distribution for \(n=10\), using \(\sigma=\sqrt{149}\) minutes. Show the substitution and include units.
  3. Why is a sample mean based on five observations more vulnerable to one unusually large value than a sample mean based on 40 observations?
  4. Describe the expected shape of the sampling distribution for \(n=15\) compared with the population’s shape. Avoid claiming exact Normality.
  5. If a simulation gives a slightly different histogram each time, does that contradict the expected pattern? Explain.