One Observation or a Sample Mean?
Suppose measurements come from a population with a known mean and standard deviation. A question about one measurement is not the same as a question about the average of several measurements. Even when both questions use the same cutoff, the probabilities can be very different.
Let \(X\) represent the value of one randomly selected individual observation, and let \(\bar{x}\) represent the mean of a random sample of a specified size. The event \(X>c\) concerns one value; the event \(\bar{x}>c\) concerns an average. As covered in “Sampling Distribution of a Sample Mean Defined,” these are different random variables with different distributions.
If the population is Normal, the distribution of an individual \(X\) is Normal with mean \(\mu\) and standard deviation \(\sigma\). The sampling distribution of \(\bar{x}\) is also exactly Normal when observations are independent, with the same center \(\mu\) but a smaller standard deviation, \(\sigma/\sqrt{n}\). Thus, the two calculations may use the same mean and cutoff but not the same standard deviation.
Standardizing makes the distinction clear. The \(z\)-score for one observation uses \(\sigma\); the \(z\)-score for a sample mean uses \(\sigma/\sqrt{n}\). A smaller standard deviation makes values farther from the center less likely. This means a cutoff above \(\mu\) will usually have a smaller upper-tail probability for \(\bar{x}\) than for \(X\). A cutoff below \(\mu\) will usually have a larger probability of being exceeded by \(\bar{x}\).
A Worked Comparison
The main task is to identify what the random variable represents before choosing a distribution. Then use the relevant spread and find the probability to the right of the cutoff. The following example uses one population and one cutoff to compare both questions directly.
Define whether the event concerns one observation \(X\) or the sample mean \(\bar{x}\), and identify the cutoff \(c\).
Identify the appropriate distribution and standard deviation. Check the sampling assumptions and whether the Normal model is exact or approximate.
Standardize the cutoff or use normalcdf with the correct mean and standard deviation for each random variable.
Interpret each probability in context, making clear whether it describes individual observations or sample means.
Worked Example: Individual Bottle Fill or Sample Mean?
Suppose fill volumes from a bottling process follow a Normal population with mean \(\mu=500\) milliliters and standard deviation \(\sigma=12\) milliliters. Individual fills are independent. Compare the probability that one fill exceeds 506 milliliters with the probability that the mean of 36 fills exceeds 506 milliliters.
State. Let \(X\) be the fill volume, in milliliters, of one bottle, and let \(\bar{x}\) be the mean fill volume for a random sample of 36 bottles. We want to compare \(P(X>506)\) and \(P(\bar{x}>506)\).
Plan. The population is Normal, so \(X\) is exactly Normal with mean 500 and standard deviation 12. The observations are independent, so \(\bar{x}\) is exactly Normal with mean 500 and standard deviation \(\sigma/\sqrt{n}\). These independent observations are drawn from a production process; if instead the sample were selected without replacement from a finite batch, we would check the 10% condition \(n\leq0.10N\).
Do: one bottle. Use the population standard deviation for an individual fill. Its standardized cutoff is:
Using normalcdf with lower bound 506, a large upper bound, mean 500, and standard deviation 12 gives:
As a check, the standard Normal upper-tail probability above \(z=0.5\) is \(1-0.6915=0.3085\), rounded to four decimal places.
Do: mean of 36 bottles. First find the standard deviation of the sample mean:
The variance check gives \(12^2/36=144/36=4\) square milliliters, and \(\sqrt{4}=2\) milliliters. Now standardize the same cutoff using the sampling-distribution standard deviation:
Using normalcdf with mean 500 and standard deviation 2 gives:
The standard Normal upper-tail probability above \(z=3\) is about 0.00135, which rounds to 0.0013 to four decimal places.
Conclude. Under the stated model, about 30.85% of individual bottle fills exceed 506 milliliters, while only about 0.13% of sample means from samples of 36 bottles exceed 506 milliliters. The cutoff is three standard deviations above the center of the sample-mean distribution but only half a standard deviation above the center of the individual-value distribution.
How the Cutoff’s Location Matters
The bottling example used a cutoff above the population mean. Because \(\bar{x}\) has a smaller standard deviation than \(X\), sample means are less likely to reach a high cutoff. If a cutoff is below the population mean, the comparison goes in the opposite direction: sample means are more likely than individual values to be above that low cutoff.
For a symmetric Normal population, a cutoff equal to \(\mu\) gives an upper-tail probability of 0.5 for both \(X\) and \(\bar{x}\). This is because both Normal distributions are centered at \(\mu\). The probabilities differ when the cutoff is away from the center, and how much they differ depends on the sample size as well as the cutoff.
| Cutoff position | Typical comparison for \(P(X>c)\) and \(P(\bar{x}>c)\) |
|---|---|
| \(c>\mu\) | \(P(\bar{x}>c)\) is smaller because the sample-mean distribution is less spread out. |
| \(c=\mu\) | For symmetric Normal distributions, both probabilities are 0.5. |
| \(c<\mu\) | \(P(\bar{x}>c)\) is larger because most sample means cluster near \(\mu\), above the low cutoff. |
Worked Example: A Less Extreme Cutoff for Bottle Fills
Use the same bottling population, with \(\mu=500\) milliliters and \(\sigma=12\) milliliters, but change the cutoff to 502 milliliters. Compare the chance that one fill and the mean of 36 fills exceed this cutoff.
State. \(X\) represents one fill, and \(\bar{x}\) represents the mean fill for a sample of 36 independent bottles. We seek \(P(X>502)\) and \(P(\bar{x}>502)\).
Plan. The Normal population gives an exact Normal model for \(X\) and, with independent observations, an exact Normal model for \(\bar{x}\). Use standard deviations 12 milliliters and 2 milliliters, respectively. The sample-mean standard deviation is \(12/\sqrt{36}=2\) milliliters.
Do: one bottle. The individual-value \(z\)-score and upper-tail probability are:
Therefore, \(\operatorname{normalcdf}(502,\infty,500,12)\approx0.4338\). A check using the standard Normal distribution gives \(1-\Phi(0.1667)\approx1-0.5662=0.4338\), rounded to four decimal places.
Do: mean of 36 bottles. The sample-mean \(z\)-score and upper-tail probability are:
Thus, \(\operatorname{normalcdf}(502,\infty,500,2)\approx0.1587\). The check is \(1-\Phi(1)=1-0.8413=0.1587\), rounded to four decimal places.
Conclude. About 43.38% of individual fills exceed 502 milliliters, compared with about 15.87% of means from samples of 36 fills. The cutoff is above the center in both distributions, but it is farther above the center when measured in standard deviations for \(\bar{x}\).
Worked Example: A Cutoff Below the Population Mean
Suppose the time a device takes to complete a startup sequence follows a Normal population with mean \(\mu=40\) seconds and standard deviation \(\sigma=8\) seconds. Startup times are independent. Compare the chance that one startup takes more than 38 seconds with the chance that the mean startup time for 16 sequences exceeds 38 seconds.
State. Let \(X\) be the startup time, in seconds, for one sequence, and let \(\bar{x}\) be the mean startup time for 16 independent sequences. The requested probabilities are \(P(X>38)\) and \(P(\bar{x}>38)\).
Plan. The population is Normal, so the model for \(X\) is exact. Independence makes the sampling distribution of \(\bar{x}\) exactly Normal as well. Its standard deviation is \(\sigma/\sqrt{n}\). If these 16 sequences were sampled without replacement from a finite collection, we would also check that \(16\leq0.10N\).
Do: one startup. The cutoff is below the mean. For one observation:
The probability above this cutoff is \(\operatorname{normalcdf}(38,\infty,40,8)\approx0.5987\). Equivalently, \(1-\Phi(-0.25)=1-0.4013=0.5987\), rounded to four decimal places.
Do: mean of 16 startups. First calculate and verify the standard deviation of the sample mean:
The variance check is \(8^2/16=64/16=4\) square seconds, whose square root is 2 seconds. The standardized cutoff and upper-tail probability are:
Thus, \(\operatorname{normalcdf}(38,\infty,40,2)\approx0.8413\), or \(1-\Phi(-1)=1-0.1587=0.8413\), rounded to four decimal places.
Conclude. About 59.87% of individual startup times exceed 38 seconds, while about 84.13% of means from samples of 16 sequences exceed 38 seconds. Since 38 seconds is below the population mean, the tighter sample-mean distribution places more of its area above this cutoff.
What the CLT Does—and Does Not—Tell You
The examples used a Normal population, so both the individual-value and sample-mean models were exactly Normal. When a population is not Normal, take care: the Central Limit Theorem concerns the sampling distribution of \(\bar{x}\), not the distribution of individual observations \(X\). As discussed in “Central Limit Theorem Explained” and “Is \(n=30\) Large Enough for the CLT,” the population shape and sample size determine whether a Normal approximation for sample means is reasonable.
A large sample does not make individual observations Normal. To calculate \(P(X>c)\), use a Normal model only if the population itself is Normal or if a suitable model for individual values is otherwise justified. A CLT-based approximation might support calculating \(P(\bar{x}>c)\), but it does not support calculating \(P(X>c)\).
For random samples taken without replacement, check the 10% condition \(n\leq0.10N\) before treating observations as independent. For independent observations by design, state that independence is assumed. These checks matter for the sample-mean calculation because its sampling distribution depends on the sampling process.
Common Mistakes and AP Exam Tip
- Using \(\sigma/\sqrt{n}\) for an individual value: \(X\) describes one observation, so use the population standard deviation \(\sigma\). The standard deviation \(\sigma/\sqrt{n}\) belongs to \(\bar{x}\).
- Using \(\sigma\) for a sample mean: Averages vary less than individual observations under the stated independence conditions. For \(\bar{x}\), use \(\sigma/\sqrt{n}\), not \(\sigma\).
- Mixing up the events: \(P(X>c)\) describes an individual measurement exceeding \(c\). \(P(\bar{x}>c)\) describes a sample average from samples of a specified size exceeding \(c\). Name the random variable in the conclusion.
- Using the CLT for individual values: The CLT helps describe the sampling distribution of a mean; it does not turn the population distribution of \(X\) into a Normal distribution.
- Assuming the sample-mean probability is always smaller: That is generally true for a cutoff above \(\mu\), but not for a cutoff below \(\mu\). Compare the cutoff’s location with the center before predicting the result.
- Reporting a calculator value without explaining its inputs: Show the mean and standard deviation used for each distribution, and include the context and units in the interpretation.
For full-credit communication, define \(X\) and \(\bar{x}\), state the relevant population or sampling-distribution parameters, justify the model, and interpret each probability with its subject and context. Make clear that the same cutoff does not make the two probability questions equivalent.
Check Your Understanding
For each question, identify the random variable, select the appropriate standard deviation, and explain what the probability describes.
- A Normal population has mean 72 centimeters and standard deviation 10 centimeters. Compare \(P(X>77)\) with \(P(\bar{x}>77)\) for independent samples of size 25. Which probability is smaller, and why?
- A Normal population has mean 18 minutes and standard deviation 6 minutes. For independent samples of size 9, find \(P(X>15)\) and \(P(\bar{x}>15)\). Interpret both probabilities.
- Individual observations come from a strongly skewed population. A student says the CLT makes \(X\) approximately Normal for a sample of size 100. Explain the error.
- For a symmetric Normal population, what are \(P(X>\mu)\) and \(P(\bar{x}>\mu)\)? Explain why they are equal.
- A student calculates \(P(\bar{x}>c)\) using \(\sigma\) instead of \(\sigma/\sqrt{n}\). Describe how this confuses the individual-value distribution with the sampling distribution.