Tutorials › AP Statistics › Mean of the Sampling Distribution of x-bar

Sampling distributions for means · Tutorial 602 of 1000

Mean of the Sampling Distribution of x-bar

See why sample means center on the population mean across repeated random sampling, even though individual sample means can vary.

Intermediate 9 min read

What You'll Learn

  • Calculate the mean of a sampling distribution from its possible values and probabilities
  • Verify that the mean of the sampling distribution of x-bar equals the population mean
  • Explain what it means for x-bar to be an unbiased estimator of mu
  • Check the result for samples taken with and without replacement
  • Distinguish unbiasedness from a guarantee that any one sample mean equals mu

The Center of the Sampling Distribution

In the previous tutorial, “Sampling Distribution of a Sample Mean Defined,” we built a probability distribution from the possible values of \(\bar{x}\). That distribution describes how sample means vary across all samples of a fixed size selected by a specified method. A natural next question is: where is this distribution centered?

For a random sample from a population with mean \(\mu\), the mean of the sampling distribution of \(\bar{x}\) is \(\mu\). This is true whether a particular sample mean falls above or below the population mean. Some possible samples may produce high means and others low means, but the probabilities balance so that the average of all possible sample means is the population mean.

Key result: For random samples of a fixed size from a population with mean \(\mu\), the mean of the sampling distribution of \(\bar{x}\) is \(\mu\). In symbols, \(\mu_{\bar{x}}=\mu\).

Here, \(\mu_{\bar{x}}\) means the mean of the sampling distribution of the sample mean, not the mean of one particular sample. We can calculate it just as we calculate the expected value of any discrete probability distribution: multiply each possible value by its probability, then add the products.

$$ \mu_{\bar{x}}=\sum \bar{x}\,P(\bar{x}) $$

The notation \(\bar{x}\) represents a possible sample mean in the distribution, and \(P(\bar{x})\) is the probability of getting that value. The population mean \(\mu\), by contrast, describes the population’s individual values. The result \(\mu_{\bar{x}}=\mu\) says the sampling distribution is centered at the population mean; it does not say that every sample mean equals \(\mu\).

Why the Two Means Are Equal

For a small population and sample size, we can check the result by listing every possible sample, as in the previous tutorial. When samples are selected by simple random sampling, every population member has the same chance of being included. That equal chance keeps the sample mean from systematically favoring some population values over others.

The idea also works for larger populations, where listing all samples would be impractical. Suppose the population has \(N\) values, with mean \(\mu\), and a sample of size \(n\) is selected. Without replacement, each population value has probability \(n/N\) of appearing in the sample, so the expected total of the selected values is \(n\mu\). With replacement, each of the \(n\) draws has expected value \(\mu\), so the expected total is also \(n\mu\). Dividing that total by \(n\), as we do to get a sample mean, gives \(\mu\).

$$ \text{Mean of the sample total}=n\mu \qquad\Longrightarrow\qquad \mu_{\bar{x}}=\frac{n\mu}{n}=\mu $$

This result applies to simple random samples of a fixed size, whether the sampling is with or without replacement. With replacement, each draw has the population distribution. Without replacement, the selected values are dependent, but each population member still has the same chance of inclusion in a simple random sample. The result concerns the center of the sampling distribution; the spread of that distribution is a separate question.

A statistic whose sampling distribution has mean equal to the population parameter it estimates is called an unbiased estimator of that parameter. Thus, for random samples taken as described here, \(\bar{x}\) is an unbiased estimator of \(\mu\).

Definition: An estimator is unbiased for a population parameter if the mean of its sampling distribution equals that parameter. Since \(\mu_{\bar{x}}=\mu\), the sample mean \(\bar{x}\) is an unbiased estimator of the population mean \(\mu\).

“Unbiased” does not mean that a sample mean is always correct or that it will equal \(\mu\) in a particular sample. It means that, over the full set of possible samples under the stated random sampling method, the average of the sample means equals \(\mu\). A single sample can still produce an estimate that is noticeably too high or too low.

Verify the Result When \(\mu=50\)

The following example uses the population values \(30,50,70\), whose mean is \(\mu=50\). Drawing two values with replacement gives nine equally likely ordered samples. Listing them makes it possible to see exactly how the sampling distribution is centered.

Worked Example: The Sampling Distribution Centers at 50

A population consists of \(30,50,70\). Two values are selected with replacement, with each value equally likely on each selection. Find the sampling distribution of \(\bar{x}\), calculate its mean, and explain whether \(\bar{x}\) is unbiased for \(\mu\).

Step 1: Find the population mean. The mean of the three population values is:

$$ \mu=\frac{30+50+70}{3}=\frac{150}{3}=50 $$

Step 2: List the equally likely outcomes and their sample means. Because the sample is taken with replacement, there are \(3\times3=9\) ordered samples. For example, \((30,50)\) and \((50,30)\) are separate outcomes, each with probability \(1/9\). Grouping outcomes that produce the same mean gives:

Possible value of \(\bar{x}\)Number of outcomesProbability
301\(1/9\)
402\(2/9\)
503\(3/9\)
602\(2/9\)
701\(1/9\)

For instance, a mean of 50 results from \((30,70)\), \((50,50)\), or \((70,30)\), so its probability is \(3/9\). The probabilities sum to \((1+2+3+2+1)/9=9/9=1\).

Step 3: Calculate the mean of the sampling distribution. Multiply each possible sample mean by its probability and add:

$$ \begin{aligned} \mu_{\bar{x}} &=30\left(\frac{1}{9}\right)+40\left(\frac{2}{9}\right)+50\left(\frac{3}{9}\right)+60\left(\frac{2}{9}\right)+70\left(\frac{1}{9}\right)\\ &=\frac{30+80+150+120+70}{9}\\ &=\frac{450}{9}=50 \end{aligned} $$

As a check, the deviations from 50 cancel when weighted by their probabilities: \((-20)(1/9)+(-10)(2/9)+0(3/9)+(10)(2/9)+(20)(1/9)=0\). Thus the distribution’s mean is 50, matching the population mean. The sample mean is unbiased for \(\mu\) in this setting.

Unbiasedness Does Not Require Replacement

The result is not limited to sampling with replacement. In a simple random sample without replacement, the possible samples and their probabilities differ, but the mean of the sampling distribution still equals the population mean. Here is a small example that verifies this by enumeration.

Worked Example: Sampling Without Replacement

A population consists of \(40,50,60,70\). A simple random sample of size 2 is selected without replacement. Find the population mean and the mean of the sampling distribution of \(\bar{x}\).

The population mean is \(\mu=(40+50+60+70)/4=220/4=55\). There are six equally likely pairs. Their sample means are:

SampleSample mean \(\bar{x}\)
\(\{40,50\}\)45
\(\{40,60\}\)50
\(\{40,70\}\)55
\(\{50,60\}\)55
\(\{50,70\}\)60
\(\{60,70\}\)65

Each pair has probability \(1/6\), so each listed sample mean is an equally likely outcome except that 55 occurs for two pairs. The mean of the sampling distribution is:

$$ \mu_{\bar{x}}=\frac{45+50+55+55+60+65}{6} =\frac{330}{6}=55 $$

This agrees with the population mean. One way to check the arithmetic is to subtract 55 from each of the six sample means: the deviations are \(-10,-5,0,0,5,10\), which sum to 0. The sample mean is therefore unbiased for \(\mu\) under this simple random sampling method, even though some possible sample means are below 55 and others are above it.

Use the Result Without Listing Every Sample

For a larger population, a full list of possible samples may be too long to construct. The equality \(\mu_{\bar{x}}=\mu\) lets us identify the center of the sampling distribution directly, as long as the sample is selected randomly in a way that gives the population values equal representation. We can also see the reasoning by considering the sample positions.

For sampling with replacement, each position in the sample has the population distribution. Its expected value is therefore \(\mu\). Since \(\bar{x}\) is the average of the values in those positions, the mean of \(\bar{x}\) is the average of their means, also \(\mu\). For simple random sampling without replacement, each position likewise has the population mean across all possible samples, so averaging the positions again gives \(\mu\).

Worked Example: Identify the Center Without Enumerating Samples

A community garden records four possible harvest weights for a plot: \(20,40,60,80\) pounds. A random sample of three plots is taken with replacement. Find the mean of the sampling distribution of \(\bar{x}\), and explain what the result says about using \(\bar{x}\) to estimate the population mean.

Step 1: Calculate the population mean. The mean weight is:

$$ \mu=\frac{20+40+60+80}{4}=\frac{200}{4}=50\text{ pounds} $$

Step 2: Use the sampling-distribution result. Each draw has expected value 50 pounds. A sample of three draws has a sample mean, so the mean of its sampling distribution is:

$$ \mu_{\bar{x}}=\frac{50+50+50}{3}=50\text{ pounds} $$

There are \(4^3=64\) ordered samples, but listing all of them is unnecessary to find the center. As a check on why individual results can differ, the sample \((20,20,20)\) has mean 20 pounds, while \((80,80,80)\) has mean 80 pounds. Those are possible sample means, yet across all equally likely samples the distribution has mean 50 pounds. Therefore, \(\bar{x}\) is unbiased for the population mean of 50 pounds; it is not guaranteed to equal 50 for any one sample.

What Unbiasedness Does—and Does Not—Tell You

Unbiasedness describes the location of a sampling distribution, not how tightly its values cluster. Two sampling distributions can both be centered at \(\mu\) while one has more spread than the other. The mean alone does not tell us how close a particular sample mean is likely to be to \(\mu\); the next tutorial considers the standard deviation of the sample mean.

It is also important to connect the conclusion to the sampling method. The result applies to random sampling that treats population members fairly, such as a simple random sample. If a method systematically gives some population values a greater chance of selection than others, the sample mean may no longer be centered at the population mean. Do not claim unbiasedness without considering how the samples are selected.

Common Mistakes and AP Exam Tip

  • Confusing \(\mu\) with \(\mu_{\bar{x}}\): \(\mu\) is the mean of individual population values. \(\mu_{\bar{x}}\) is the mean of the possible sample means. State which distribution you are describing.
  • Assuming every sample mean equals the population mean: Unbiasedness concerns the average over all possible samples, not the result for one sample. A full-credit answer says that the sampling distribution is centered at \(\mu\), not that every \(\bar{x}\) equals \(\mu\).
  • Ignoring probabilities when calculating a distribution’s mean: If possible sample means are not equally likely, multiply each value by its probability. Do not simply average the distinct values unless they really are equally likely outcomes.
  • Counting unordered and ordered samples inconsistently: With replacement, ordered draws such as \((30,50)\) and \((50,30)\) are distinct outcomes. With a simple random sample without replacement, list each pair once when order does not matter.
  • Calling a statistic unbiased without naming the parameter: Say that \(\bar{x}\) is an unbiased estimator of the population mean \(\mu\), because \(\mu_{\bar{x}}=\mu\).
  • Overstating what the result guarantees: The result does not promise that one observed sample mean is close to \(\mu\). It identifies the sampling distribution’s mean; its spread is a different feature.

For a full-credit explanation, identify the population mean, state that the sample is selected by an appropriate random sampling method, and connect the equality \(\mu_{\bar{x}}=\mu\) to unbiasedness. If the question gives a small sampling distribution, show the probability-weighted calculation and check the arithmetic.

Key takeaway: For random samples of a fixed size selected by simple random sampling, the mean of the sampling distribution of \(\bar{x}\) equals the population mean: \(\mu_{\bar{x}}=\mu\). This makes \(\bar{x}\) an unbiased estimator of \(\mu\), but it does not guarantee that any particular sample mean equals the population mean.

Check Your Understanding

Answer each question using the distinction between the population mean and the mean of the sampling distribution.

  1. A population has mean \(\mu=50\). What is the mean of the sampling distribution of \(\bar{x}\) for a simple random sample of fixed size?
  2. In a small population, possible sample means are 40, 50, and 60, with probabilities \(1/4,1/2,\) and \(1/4\). Calculate \(\mu_{\bar{x}}\).
  3. Explain in context what it means to say that \(\bar{x}\) is an unbiased estimator of \(\mu\).
  4. Does \(\mu_{\bar{x}}=\mu\) mean every sample mean equals the population mean? Explain.
  5. Why is it important to consider the sampling method before stating that \(\bar{x}\) is unbiased?