Tutorials › AP Statistics › Sampling Distribution of a Mean in Context Problems

Sampling distributions for means · Tutorial 619 of 1000

Sampling Distribution of a Mean in Context Problems

Practice translating real-world questions into a distribution for the sample mean, checking its conditions, and explaining probability results in context.

Intermediate 9 min read

What You'll Learn

  • Identify the random quantity and describe the sampling distribution of the sample mean in context.
  • State the center, standard deviation, and units of the sampling distribution.
  • Check random selection, independence, the 10% condition, and whether a Normal model is appropriate.
  • Use a finite population correction when sampling without replacement and the 10% condition is not met.
  • Calculate and interpret probabilities about fill weights, commute times, and other sample means.

Turn the Story Into a Distribution for the Sample Mean

A contextual problem about fill weights, commute times, or service times can mention several quantities at once: individual measurements, a population, and a sample. Start by identifying what varies from sample to sample. If the question asks about the average for a sample of a particular size, the random quantity is the sample mean \(\bar{x}\), and the relevant model is its sampling distribution.

As explained in “Sampling Distribution of a Sample Mean Defined,” the sampling distribution describes the possible values of \(\bar{x}\) from samples of the same size selected in the same way. Earlier tutorials established that its center is \(\mu\), and that for independent observations its standard deviation is \(\sigma/\sqrt{n}\). In a context problem, a complete description names the quantity, center, spread, shape, and units—and explains why the model is appropriate.

Context checklist: Identify the random quantity (\(\bar{x}\)); give its center and standard deviation with units; describe its shape as Normal or approximately Normal; and check the sampling design and conditions that support the model.

Build the Distribution Before Finding a Probability

A useful first step is to translate the problem’s wording into statistical language. Record the population mean \(\mu\), population standard deviation \(\sigma\), sample size \(n\), and whether the sample is selected with or without replacement. Keep the units attached: if individual fill weights are measured in grams, then both \(\mu\) and \(\sigma_{\bar{x}}\) are measured in grams.

For independent observations, \(\mu_{\bar{x}}=\mu\) and \(\sigma_{\bar{x}}=\sigma/\sqrt{n}\). If the observations come from a simple random sample without replacement from a finite population of size \(N\), consider the 10% condition, as covered in “The 10% Condition for Sample Means.” When \(n<0.10N\), treating observations as independent is reasonable for the standard deviation calculation. If that condition is not met, the exact standard deviation for a simple random sample without replacement uses the finite population correction.

$$ \sigma_{\bar{x}}= \frac{\sigma}{\sqrt{n}} \sqrt{\frac{N-n}{N-1}}. $$

This formula assumes that \(\sigma\) is the standard deviation of the \(N\) values in the finite population. The correction factor accounts for the fact that sampling without replacement makes observations dependent: once a value is selected, it cannot be selected again. It reduces the standard deviation compared with \(\sigma/\sqrt{n}\). The 10% condition is not met in the examples where we use this correction; the correction, rather than an independence approximation, accounts for the finite-population sampling.

Next decide whether a Normal model for \(\bar{x}\) is reasonable. If individual observations are independent and come from a Normal population, \(\bar{x}\) is Normal for any sample size, as explained in “Sampling Distribution When the Population Is Normal.” If the population is not Normal, use the CLT with care: a sufficiently large random sample can make the sampling distribution approximately Normal, but the required sample size depends on the population’s shape. A sample size alone does not establish that a Normal model is appropriate.

1
Define the quantity.
Say whether the question concerns one observation \(X\) or a sample mean \(\bar{x}\), and identify the sample size.
2
Check the design.
Establish whether the sample is random and whether observations can reasonably be treated as independent. For a simple random sample without replacement, check the 10% condition.
3
Describe the distribution.
Give its center, standard deviation, shape, and units. If the 10% condition is not met for a finite population, consider the finite population correction.
4
Calculate and interpret.
Use the distribution of \(\bar{x}\) to find the requested probability, then state what it means for sample means in the given context.

Worked Examples: Fill Weights and Commute Times

Worked Example: A Random Sample of Bottles From a Finite Lot

A hypothetical lot contains 300 bottles. The mean fill weight is 500 grams, and the standard deviation of the lot’s fill weights is 4.8 grams. The lot’s weights are approximately Normal. A simple random sample of 45 bottles is drawn without replacement. Estimate the probability that the sample mean fill weight is less than 498.7 grams.

State. The random quantity is \(\bar{x}\), the mean fill weight for a random sample of 45 bottles. The question is about sample means, not the probability that one bottle weighs less than 498.7 grams.

Plan. The sample is a simple random sample, so the selection is random. Check the 10% condition: \(45<0.10(300)=30\) is false. The sample is more than 10% of the lot, so the independent-observations formula \(\sigma/\sqrt{n}\) alone is not appropriate. Use the finite population correction for the standard deviation. The lot’s weights are approximately Normal, so a Normal model for the sample mean is reasonable here. We will use that model to estimate the requested probability.

Do. The sampling distribution is centered at the lot mean, 500 grams. Its standard deviation, including the finite population correction, is

$$ \sigma_{\bar{x}} =\frac{4.8}{\sqrt{45}}\sqrt{\frac{300-45}{300-1}} =\frac{4.8}{\sqrt{45}}\sqrt{\frac{255}{299}} \approx 0.660799\text{ grams}. $$

As a check, \(4.8/\sqrt{45}\approx0.715542\) grams, and the correction factor \(\sqrt{255/299}\approx0.923495\). Their product is approximately \(0.660799\) grams. Standardize 498.7 grams using this spread:

$$ z=\frac{498.7-500}{0.660799}\approx-1.9673. $$

Thus, using a Normal model, \(P(\bar{x}<498.7)\approx P(Z<-1.9673)\approx0.0246\), rounded to four decimal places.

Conclude. Under the stated sampling design and Normal approximation, the probability that a random sample of 45 bottles from this lot has a mean fill weight below 498.7 grams is about 0.0246. The calculation uses the standard deviation of the sample mean with a finite population correction, not the standard deviation of individual bottle weights.

Worked Example: A Sample of Commute Times From a Large Population

In a hypothetical city, commute times are right-skewed, with population mean \(\mu=28\) minutes and population standard deviation \(\sigma=12\) minutes. The distribution is not extremely skewed and has no unusually extreme outliers. A simple random sample of 100 commuters is selected from 20,000 commuters. Estimate the probability that the sample mean commute time exceeds 30.4 minutes.

State. The random quantity is \(\bar{x}\), the mean commute time for a random sample of 100 commuters.

Plan. The sample is random. Check the 10% condition: \(100<0.10(20{,}000)=2{,}000\), so it is satisfied; treating observations as independent is reasonable. The population is right-skewed, not Normal, so the sampling distribution is not guaranteed to be exactly Normal. However, \(n=100\), and the population is described as not extremely skewed and without unusually extreme outliers. A Normal approximation for the sampling distribution is reasonable by the CLT. We do not claim that individual commute times are Normal.

Do. The sampling distribution has mean 28 minutes and standard deviation

$$ \mu_{\bar{x}}=\mu=28\text{ minutes}, \qquad \sigma_{\bar{x}}=\frac{12}{\sqrt{100}}=1.2\text{ minutes}. $$

Standardize the cutoff:

$$ z=\frac{30.4-28}{1.2}=2.00. $$

Using the Normal approximation, \(P(\bar{x}>30.4)\approx P(Z>2.00)=0.0228\), rounded to four decimal places. As a check, 30.4 is two standard deviations of the sampling distribution above its center, since \(28+2(1.2)=30.4\).

Conclude. Under the stated sampling conditions and the CLT approximation, the probability that the mean commute time in a random sample of 100 commuters exceeds 30.4 minutes is about 0.0228. This is a probability about sample means, not about the fraction of individual commuters whose commute exceeds 30.4 minutes.

Worked Example: A Mean From a Normally Distributed Process

A hypothetical production process has Normal measurement values with mean \(\mu=42\) seconds and standard deviation \(\sigma=6\) seconds. Measurements from the ongoing process can be treated as independent. A sample of 16 measurements is taken. Find the probability that the sample mean is between 40.5 and 43.5 seconds.

State. The random quantity is \(\bar{x}\), the mean of 16 process measurements. The question asks for the probability that this sample mean falls within the stated interval.

Plan. The observations are independent and come from a Normal population. Therefore, the sampling distribution of \(\bar{x}\) is exactly Normal, even though the sample size is only 16. This is an ongoing process model, not a simple random sample without replacement from a finite population, so a 10% condition is not needed.

Do. The center of the sampling distribution is 42 seconds, and its standard deviation is

$$ \mu_{\bar{x}}=42\text{ seconds}, \qquad \sigma_{\bar{x}}=\frac{6}{\sqrt{16}}=1.5\text{ seconds}. $$

The lower and upper endpoints are each 1.5 seconds from the center:

$$ z_{\text{lower}}=\frac{40.5-42}{1.5}=-1, \qquad z_{\text{upper}}=\frac{43.5-42}{1.5}=1. $$

Therefore, \(P(40.5<\bar{x}<43.5)=P(-1<Z<1)\approx0.6827\), rounded to four decimal places. The interval is one standard deviation below to one standard deviation above the mean of the sampling distribution.

Conclude. The probability that the mean of 16 independent measurements from this Normal process is between 40.5 and 43.5 seconds is about 0.6827. The Normal model here is exact because the population is Normal and the observations are independent, not because the sample size is large.

Common Mistakes and AP Exam Tip

  • Describing individual observations instead of sample means. If the question asks for an average from a sample, define \(\bar{x}\). Do not use \(\sigma\) as the spread of \(\bar{x}\); use the appropriate standard deviation of the sampling distribution.
  • Skipping the sampling design. A formula does not make a sample random or observations independent. State how the sample was selected, and check the 10% condition when sampling without replacement from a finite population.
  • Applying the independence formula when the 10% condition fails. For a simple random sample without replacement when \(n\) is not less than \(0.10N\), use the finite population correction for the standard deviation rather than ignoring the dependence.
  • Claiming the CLT makes the distribution exactly Normal. The CLT supports an approximation for sample means as sample size increases. It does not make individual observations Normal or guarantee a good approximation for every population and sample size.
  • Leaving units or context out of the description. State the center and standard deviation in the original units, then interpret the probability as a statement about sample means in the situation.
  • Reporting a probability without naming the model. Say whether the model is exactly Normal or approximately Normal, and give the reason: a Normal population with independent observations, or a reasonable CLT approximation under the conditions.

For full credit, a contextual response should make the chain of reasoning visible: define \(\bar{x}\), check the sampling design and conditions, describe its center, spread, and shape, show the calculation, and interpret the result in context. “Approximately Normal” is not a substitute for explaining why the approximation is reasonable.

Key takeaway: For a contextual sample-mean problem, define \(\bar{x}\) and describe its distribution before calculating. Give the center, standard deviation, shape, and units; check random selection and independence; and use the finite population correction when a simple random sample without replacement does not meet the 10% condition.

Check Your Understanding

For each question, identify the random quantity and explain which conditions or distribution facts matter.

  1. A simple random sample of 60 items is drawn without replacement from a lot of 1,000. Is the 10% condition met? What does that suggest about treating observations as independent?
  2. Independent samples of size 25 are taken from a Normal population with mean 18 centimeters and standard deviation 5 centimeters. State the distribution of \(\bar{x}\), including its center and standard deviation.
  3. A population of travel times is strongly right-skewed. Explain why a sample size of 8 does not automatically justify a Normal model for \(\bar{x}\).
  4. A random sample of 80 households is selected without replacement from 500 households. The 10% condition is not met. Name the adjustment to the standard deviation of \(\bar{x}\) that accounts for sampling without replacement.
  5. Explain the difference between \(P(X>c)\) and \(P(\bar{x}>c)\) in a problem about individual service times and sample mean service times.