Make Every Part of the Response Do a Job
In “Sampling Distribution of a Mean in Context Problems,” you practiced identifying \(\bar{x}\), describing its sampling distribution, and using that distribution to find a probability. This tutorial focuses on communicating that reasoning clearly. A strong free-response answer does more than list a formula or give a calculator result: it connects the sample, the model, and the conclusion to the actual situation.
A useful response usually moves in this order: identify the random quantity; justify the model’s shape; give its center and spread with units; show the probability calculation, if the model supports one; and interpret the result in context. Not every question requires a probability. If the information does not justify a Normal model or does not specify enough about the population distribution, say so instead of presenting an unsupported number.
Shape, center, and spread are separate claims and need separate support. A Normal population with independent observations gives an exactly Normal sampling distribution. For a non-Normal population, the Central Limit Theorem may support an approximately Normal sampling distribution for a sufficiently large sample, but the population’s shape matters too. The center and spread calculations do not, by themselves, prove that a Normal model is appropriate.
As in “The 10% Condition for Sample Means,” when a simple random sample is taken without replacement from a finite population, check whether \(n<0.10N\). If that condition is met, treating observations as independent is reasonable for the usual standard deviation calculation. The standard deviation is about the variation of sample means, not the variation of individual observations. Keep the units attached: if individual measurements are in seconds, the center and standard deviation of \(\bar{x}\) are also in seconds.
When writing a probability statement, make clear what event is being measured. For example, \(P(\bar{x}>c)\) concerns the chance that the mean of a random sample of a specified size exceeds \(c\). It is not the chance that an individual observation exceeds \(c\), and it does not say what will happen in every sample.
A Reliable Sequence for Free-Response Writing
Define \(\bar{x}\) in context and give the sample size. This makes clear that the response concerns sample means rather than individual measurements.
Use the population’s Normal shape or explain why a CLT approximation is reasonable. Mention random selection and independence, and check the 10% condition for sampling without replacement from a finite population.
State the mean and standard deviation of the sampling distribution in the original units. Show the substitution used to calculate the standard deviation.
Show the standardized cutoff or calculator setup, report the probability with consistent rounding, and explain what it means for samples in this context.
This sequence is a communication tool, not a requirement to repeat every fact in a long paragraph. Concise statements can earn credit when they identify the relevant quantity, give evidence for the model, and state what the result means. In particular, write “exactly Normal” only when the assumptions make that claim exact; otherwise, use “approximately Normal” and name the reason for the approximation.
Worked Responses: From Model to Context
Worked Example: Readings From a Water-Quality Sensor
A stable sensor process produces independent readings from a Normal population with mean 12.0 milligrams per liter and standard deviation 1.5 milligrams per liter. A random sample of 9 readings is taken. Find the probability that the sample mean is greater than 13.0 milligrams per liter.
State. Let \(\bar{x}\) be the mean of a random sample of 9 sensor readings, measured in milligrams per liter. We want \(P(\bar{x}>13.0)\).
Plan. The readings are a random sample from the stable process, and the problem states that they are independent. The individual readings come from a Normal population, so the sampling distribution of \(\bar{x}\) is exactly Normal for this sample size. The readings are from an ongoing process rather than a simple random sample without replacement from a stated finite population, so a 10% condition is not needed here.
Do. The sampling distribution is centered at the population mean. Its standard deviation is the population standard deviation divided by the square root of the sample size:
The standardized value of 13.0 milligrams per liter is
Thus, \(P(\bar{x}>13.0)=P(Z>2.00)\approx0.0228\), rounded to four decimal places. A calculator gives the same result using the Normal model with lower bound 13.0, upper bound a sufficiently large value, mean 12.0, and standard deviation 0.5.
Conclude. The probability that the mean of a random sample of 9 readings from this Normal sensor process exceeds 13.0 milligrams per liter is about 0.0228. This is a probability about sample means, and the Normal model is exact because the population is Normal and the readings are independent.
Worked Example: Average Loading Time for a Technology Service
A technology service has right-skewed loading times, with population mean 4.8 seconds and population standard deviation 2.0 seconds. The distribution is not extremely skewed and has no unusually extreme outliers. A random sample of 100 loading times is selected from 50,000 service sessions. Estimate the probability that the sample mean is between 4.5 and 5.1 seconds.
State. Let \(\bar{x}\) be the mean loading time for a random sample of 100 service sessions. We want \(P(4.5<\bar{x}<5.1)\), with times measured in seconds.
Plan. The sample is random. The 10% condition is met because \(100<0.10(50{,}000)=5{,}000\), so treating the observations as independent is reasonable. Individual loading times are right-skewed, so the sample mean is not exactly Normal on that basis. However, the sample size is large, and the population is described as not extremely skewed and without unusually extreme outliers. A Normal approximation for the sampling distribution is reasonable by the CLT.
Do. The center of the sampling distribution is 4.8 seconds. Its standard deviation is
The endpoints are each 0.3 seconds from the center, which is 1.5 standard deviations:
Using the approximate Normal model, \(P(4.5<\bar{x}<5.1)\approx P(-1.5<Z<1.5)\approx0.8664\), rounded to four decimal places.
Conclude. Under the stated sampling conditions and CLT approximation, about 0.8664 of random samples of 100 sessions are expected to have a mean loading time between 4.5 and 5.1 seconds. This is an approximate probability about sample means, not a statement about the percentage of individual sessions in that interval.
Worked Example: When the Population Shape Does Not Support a Probability Model
Seedling heights in a hypothetical greenhouse population of 800 plants are strongly right-skewed. The population mean height is 24 centimeters and the population standard deviation is 9 centimeters. A simple random sample of 10 plants is selected without replacement. A student is asked to estimate the probability that the sample mean height exceeds 28 centimeters.
State. Let \(\bar{x}\) be the mean height, in centimeters, for a simple random sample of 10 plants. The requested event is \(\bar{x}>28\) centimeters.
Plan. The sample is a simple random sample, so the random-selection condition is met. The 10% condition is also met: \(10<0.10(800)=80\), so treating observations as independent is reasonable for describing the standard deviation. However, the population is strongly right-skewed and the sample size is only 10. The population is not stated to be Normal, and these facts do not establish that the sampling distribution is approximately Normal. A Normal probability calculation is therefore not justified from the information given.
Do. The sampling distribution of \(\bar{x}\) is centered at 24 centimeters. Its standard deviation is
These describe the center and spread, but they do not determine the exact shape of the sampling distribution or the requested probability. The population’s mean and standard deviation alone are not enough to calculate \(P(\bar{x}>28)\) exactly. Without additional information about the population distribution or a justified approximation, we should not report a Normal-model probability.
Conclude. The sample mean has a sampling distribution centered at 24 centimeters with standard deviation about 2.846 centimeters. But the information given does not justify a Normal approximation for this sample size and strongly right-skewed population, so a reliable numerical estimate of the probability that the sample mean exceeds 28 centimeters cannot be made using the Normal model.
Common Mistakes and What a Full-Credit Response Says
- Giving a shape without a reason. “The distribution is Normal” is incomplete if the question requires justification. Say whether the population itself is Normal or explain why the CLT supports an approximate Normal model.
- Calling an approximation exact. A large sample from a non-Normal population may support an approximately Normal sampling distribution; it does not make the distribution exactly Normal. Use the word “approximately” and give the reason.
- Using the population standard deviation as the spread of \(\bar{x}\). The standard deviation of individual measurements is \(\sigma\). When the independence conditions allow it, the standard deviation of the sample mean is \(\sigma/\sqrt{n}\). Show the substitution so the reader can tell which quantity you used.
- Omitting the sampling design. A random sample and an independence check support the model. For sampling without replacement from a finite population, state whether the 10% condition is met rather than assuming it.
- Reporting a probability without its meaning. A decimal alone does not answer the contextual question. Write what type of samples the probability describes, what their mean measures, and the event or interval in the problem.
- Making a probability claim when the model is unsupported. If the population is strongly skewed and the sample is small, do not use a Normal calculation automatically. State what can be determined from the information and what cannot be justified.
A helpful final check is to read the response as if the formulas were removed. Can a reader tell what is being sampled, why the shape claim is reasonable, what the center and spread mean in the original units, and what the probability refers to? If not, add the missing connection. A complete answer need not be lengthy, but each statistical claim should have evidence and context.
Check Your Understanding
For each prompt, focus on both the statistical reasoning and how you would communicate it in a complete response.
- A Normal population has mean 36 minutes and standard deviation 8 minutes. Independent observations are sampled in groups of 16. State the shape, center, and standard deviation of the sampling distribution of \(\bar{x}\), with units and justification.
- A right-skewed population has no extreme outliers. A random sample of 81 observations is drawn from a population of 20,000. What must a response say to justify an approximate Normal model for \(\bar{x}\)?
- A simple random sample of 70 items is drawn without replacement from a population of 600. Check the 10% condition and explain what it implies for the independence assumption.
- Explain why a probability about \(X\) exceeding a cutoff is not the same statement as a probability about \(\bar{x}\) exceeding that cutoff.
- A population is strongly right-skewed, and a random sample of 7 observations is taken. The population mean and standard deviation are known. Explain why those facts alone do not necessarily justify a Normal-model probability for \(\bar{x}\).