Tutorials › AP Statistics › Standard Error Versus Standard Deviation of x-bar

Sampling distributions for means · Tutorial 613 of 1000

Standard Error Versus Standard Deviation of x-bar

Distinguish the true standard deviation of sample means from its sample-based estimate, and practice calculating each with the correct information.

Intermediate 9 min read

What You'll Learn

  • Identify when the sampling distribution’s standard deviation is calculated as sigma divided by the square root of n.
  • Calculate the estimated standard error using s divided by the square root of n.
  • Explain why the standard error estimates rather than reveals the true sampling-distribution spread.
  • Keep the units and meanings of sigma, s, the standard deviation of x-bar, and its standard error distinct.
  • Check sampling assumptions before using either calculation.

Two Similar Formulas, Two Different Roles

In “Standard Deviation of the Sample Mean,” you learned that the standard deviation of the sampling distribution of \(\bar{x}\) is \(\sigma/\sqrt{n}\), under the appropriate sampling assumptions. But in many real situations, the population standard deviation \(\sigma\) is unknown. A sample provides a way to estimate it: use the sample standard deviation \(s\), giving \(s/\sqrt{n}\).

These expressions look almost identical, but they do not mean the same thing. The first describes the actual spread of the sampling distribution when the population standard deviation is known. The second estimates that spread using sample data. Distinguishing them matters when describing how sample means vary and when carrying out inference for a population mean later in this course.

Key distinction: When \(\sigma\) is known, the standard deviation of the sampling distribution of \(\bar{x}\) is \(\sigma_{\bar{x}}=\sigma/\sqrt{n}\). When \(\sigma\) is unknown, estimate that standard deviation with the standard error \(s/\sqrt{n}\). The standard error is an estimate, not the exact sampling-distribution standard deviation.

Here, \(\sigma\) is the standard deviation of individual values in the population, and \(s\) is the standard deviation of individual values in the observed sample. Dividing either by \(\sqrt{n}\) changes the focus from the spread of individual values to the spread of sample means, or an estimate of that spread. Both results have the same measurement units as the original data.

When to Use Each Quantity

Use \(\sigma/\sqrt{n}\) when the population standard deviation is known and the sampling assumptions support the formula. For example, a process may have a well-established population standard deviation, or a problem may explicitly provide \(\sigma\). This value describes the sampling distribution of \(\bar{x}\); it is not the standard deviation of the individual observations.

Use \(s/\sqrt{n}\) when the population standard deviation is unknown but a sample is available. Since \(s\) measures the spread of the observed individual values, \(s/\sqrt{n}\) estimates how much sample means vary from sample to sample. In AP Statistics, this estimated spread is called the standard error of \(\bar{x}\). It is commonly used in procedures for inference about a population mean.

$$ \text{Known population spread: }\quad \sigma_{\bar{x}}=\frac{\sigma}{\sqrt{n}} \qquad \text{Unknown population spread: }\quad SE_{\bar{x}}=\frac{s}{\sqrt{n}} $$

The notation helps signal the difference. The symbol \(\sigma_{\bar{x}}\) names the actual standard deviation of the sampling distribution. The notation \(SE_{\bar{x}}\) emphasizes that a value calculated from \(s\) is an estimated standard error. Some settings or technology may use slightly different labels, so read the situation: is the calculation based on a known population \(\sigma\), or on an observed sample \(s\)?

Both formulas rely on an appropriate sampling process and independent observations, as discussed in “Standard Deviation of the Sample Mean” and “The 10% Condition for Sample Means.” For a simple random sample without replacement from a finite population, check the 10% condition \(n\leq0.10N\) before treating the observations as independent. If the sampling assumptions are not met, the formula may not describe the spread correctly.

Worked Examples: Calculate and Interpret the Spread

Worked Example: Known Population Standard Deviation

A sensor produces independent temperature readings from a process with a known population standard deviation of \(\sigma=9\) degrees Celsius. A quality analyst takes 36 readings and calculates their mean. Find and interpret the standard deviation of the sampling distribution of \(\bar{x}\).

State. Let \(\bar{x}\) be the mean temperature for a sample of 36 readings. The population standard deviation is known: \(\sigma=9\) degrees Celsius. We want the actual standard deviation of the sampling distribution of \(\bar{x}\).

Plan. The readings are independent, so use \(\sigma_{\bar{x}}=\sigma/\sqrt{n}\). The population spread is supplied, so there is no need to estimate it with \(s\). Because the readings are independent by design, an independence condition is stated. If instead 36 readings were sampled without replacement from a finite collection, we would check that \(36\leq0.10N\).

Do. Substitute \(\sigma=9\) degrees Celsius and \(n=36\):

$$ \sigma_{\bar{x}}=\frac{9}{\sqrt{36}}=\frac{9}{6}=1.5\text{ degrees Celsius} $$

As a check, the variance of the sample mean is \(9^2/36=81/36=2.25\) square degrees Celsius, and \(\sqrt{2.25}=1.5\) degrees Celsius.

Conclude. For repeated independent samples of 36 readings, the sample means typically vary from the process mean by about 1.5 degrees Celsius. This is the standard deviation of the sampling distribution, calculated from the known population standard deviation; it is not the spread of the 36 individual readings.

Worked Example: Estimate the Spread from a Sample

A researcher takes a simple random sample of 25 plants from a greenhouse population of 800 plants and measures their heights. The sample standard deviation is \(s=5.6\) centimeters. Estimate the standard deviation of the sampling distribution of the sample mean height.

State. Let \(\bar{x}\) be the mean height for a sample of 25 plants. The population standard deviation \(\sigma\) is unknown, but the sample standard deviation is \(s=5.6\) centimeters. We want the estimated standard error of \(\bar{x}\).

Plan. Use \(SE_{\bar{x}}=s/\sqrt{n}\), since the population spread is unknown. The sample is described as a simple random sample. The 10% condition is met because \(25\leq0.10(800)=80\); thus, treating the observations as independent is reasonable. We also assume the sample was selected as stated.

Do. Substitute \(s=5.6\) centimeters and \(n=25\):

$$ SE_{\bar{x}}=\frac{5.6}{\sqrt{25}}=\frac{5.6}{5}=1.12\text{ centimeters} $$

A variance check gives \(s^2/n=5.6^2/25=31.36/25=1.2544\) square centimeters. The square root of \(1.2544\) is \(1.12\) centimeters.

Conclude. The estimated standard error of the mean plant height is 1.12 centimeters. In repeated samples of 25 plants, sample means are expected to vary around the population mean; 1.12 centimeters estimates the typical amount of that variation. Since \(\sigma\) is unknown, this is an estimate rather than the exact standard deviation of the sampling distribution.

Worked Example: Same Sample Spread, Different Sample Sizes

Two independent simple random samples are taken from a large population of package weights. In each sample, the observed standard deviation is \(s=8\) grams. One sample has \(n=16\) packages and the other has \(n=64\). Calculate the estimated standard error of the sample mean in each case and compare them.

State. For each sample, the parameter of interest is the population mean package weight. The population standard deviation is unknown. We are comparing the estimated standard errors based on \(s=8\) grams for sample sizes 16 and 64.

Plan. Use \(s/\sqrt{n}\) for each sample because \(\sigma\) is unknown. The samples are described as random and independent, and the population is large enough that the 10% condition is assumed to be met. The matching sample standard deviations let us focus on the effect of sample size in this comparison.

Do. For the sample of 16 packages:

$$ SE_{\bar{x}}=\frac{8}{\sqrt{16}}=\frac{8}{4}=2\text{ grams} $$

For the sample of 64 packages:

$$ SE_{\bar{x}}=\frac{8}{\sqrt{64}}=\frac{8}{8}=1\text{ gram} $$

As a check, the estimated variances are \(8^2/16=64/16=4\) square grams and \(8^2/64=64/64=1\) square gram; their square roots are 2 grams and 1 gram.

Conclude. The estimated standard error is 2 grams for the sample of 16 and 1 gram for the sample of 64. With the same observed individual-value spread, quadrupling the sample size halves the estimated standard error. These are sample-based estimates: because \(\sigma\) is unknown, neither calculation gives the exact standard deviation of the sampling distribution.

Why “Standard Error” Is an Estimate

Imagine repeatedly taking samples of size \(n\) from the same population and calculating \(\bar{x}\) each time. The standard deviation of all those possible sample means is a population-level property of the sampling distribution. If \(\sigma\) is known, that standard deviation can be calculated as \(\sigma/\sqrt{n}\). If \(\sigma\) is unknown, one observed sample cannot reveal the exact spread across all possible sample means.

Instead, the observed sample gives \(s\), an estimate of the population’s individual-value spread. Dividing \(s\) by \(\sqrt{n}\) gives an estimate of the sampling-distribution spread. A different sample could produce a different \(s\), and therefore a different standard error. This sample-to-sample variation is one reason to describe \(s/\sqrt{n}\) as estimated, not exact.

The standard error is not the standard deviation of the observed sample. In the plant example, \(s=5.6\) centimeters describes the spread of individual plant heights in the sample, while \(SE_{\bar{x}}=1.12\) centimeters estimates the spread of sample means from samples of 25 plants. The two quantities answer different questions, even though one is calculated from the other.

Nor is the standard error the amount by which the sample mean is guaranteed to differ from the population mean. It describes typical variability across repeated samples under the model’s assumptions. A particular sample mean might be closer to or farther from the population mean than one standard error.

In later tutorials on inference for a mean, the estimated standard error will help measure uncertainty in \(\bar{x}\). For now, keep the central distinction clear: \(\sigma/\sqrt{n}\) is the true sampling-distribution standard deviation when \(\sigma\) is known; \(s/\sqrt{n}\) is a data-based estimate used when \(\sigma\) is unknown.

Common Mistakes and AP Exam Tip

  • Calling \(s/\sqrt{n}\) the exact standard deviation: Because \(s\) comes from a sample, this quantity estimates the sampling-distribution spread. Say “estimated standard error” unless \(\sigma\) is known and the exact formula is being used.
  • Using \(s\) by itself for sample means: \(s\) describes the spread of individual observations in the sample. To estimate the spread of sample means, divide by \(\sqrt{n}\).
  • Using \(\sigma/\sqrt{n}\) when \(\sigma\) is not known: Do not treat a sample standard deviation as though it were the known population standard deviation. Use \(s/\sqrt{n}\) and identify it as an estimate.
  • Leaving off units: If individual measurements are in centimeters, both \(\sigma/\sqrt{n}\) and \(s/\sqrt{n}\) are in centimeters—not square centimeters. Squared units belong to variances.
  • Skipping sampling conditions: State that observations are independent when appropriate. For sampling without replacement from a finite population, check the 10% condition before using the independence-based formula.

For full-credit communication, identify whether the population standard deviation is known, name the quantity being calculated, show the substitution, and interpret the result in context with units. If the calculation uses \(s\), explicitly call the result an estimated standard error. Avoid implying it is the exact distance between a sample mean and the population mean.

Key takeaway: Use \(\sigma/\sqrt{n}\) for the actual standard deviation of the sampling distribution when the population standard deviation is known. When \(\sigma\) is unknown, use \(s/\sqrt{n}\) as an estimate and call it the standard error. Both describe the spread of sample means, not individual observations.

Check Your Understanding

For each situation, decide whether to use \(\sigma/\sqrt{n}\) or \(s/\sqrt{n}\), and explain what the result represents.

  1. A process has a known population standard deviation of 14 millimeters. Find the standard deviation of \(\bar{x}\) for independent samples of size 49.
  2. A random sample of 36 trees has sample standard deviation \(s=3\) meters. The trees were sampled without replacement from a population of 500. Check the 10% condition and find the estimated standard error of the mean height.
  3. A student reports \(s=10\) seconds as the standard error for a mean based on a sample of size 25. Explain the mistake and give the correct estimated standard error.
  4. In a sample of 20 devices, \(s=6\) minutes. Explain what \(s\) describes and what \(s/\sqrt{20}\) estimates.
  5. Why can two different samples of the same size from the same population produce different estimated standard errors?