Tutorials › AP Statistics › Comparing Individual Values and Sample Means

Sampling distributions for means · Tutorial 612 of 1000

Comparing Individual Values and Sample Means

Learn why an individual value and a sample mean from the same population can have very different probabilities of exceeding the same cutoff.

Intermediate 10 min read

What You'll Learn

  • Distinguish the random variables \(X\) and \(\bar{x}\) and the events they describe.
  • Choose the correct standard deviation for an individual observation or a sample mean.
  • Calculate and compare upper-tail probabilities using the same cutoff.
  • Explain how a sample mean’s smaller spread affects probabilities above or below the population mean.
  • Identify when a Normal model is exact for individual values and for sample means.
  • Avoid using the CLT to justify a Normal model for individual observations.

One Observation or a Sample Mean?

Suppose measurements come from a population with a known mean and standard deviation. A question about one measurement is not the same as a question about the average of several measurements. Even when both questions use the same cutoff, the probabilities can be very different.

Let \(X\) represent the value of one randomly selected individual observation, and let \(\bar{x}\) represent the mean of a random sample of a specified size. The event \(X>c\) concerns one value; the event \(\bar{x}>c\) concerns an average. As covered in “Sampling Distribution of a Sample Mean Defined,” these are different random variables with different distributions.

Key distinction: To calculate \(P(X>c)\), use the population distribution of individual observations. To calculate \(P(\bar{x}>c)\), use the sampling distribution of sample means. For independent observations, that sampling distribution has mean \(\mu\) and standard deviation \(\sigma/\sqrt{n}\), as covered in “Standard Deviation of the Sample Mean.”

If the population is Normal, the distribution of an individual \(X\) is Normal with mean \(\mu\) and standard deviation \(\sigma\). The sampling distribution of \(\bar{x}\) is also exactly Normal when observations are independent, with the same center \(\mu\) but a smaller standard deviation, \(\sigma/\sqrt{n}\). Thus, the two calculations may use the same mean and cutoff but not the same standard deviation.

Standardizing makes the distinction clear. The \(z\)-score for one observation uses \(\sigma\); the \(z\)-score for a sample mean uses \(\sigma/\sqrt{n}\). A smaller standard deviation makes values farther from the center less likely. This means a cutoff above \(\mu\) will usually have a smaller upper-tail probability for \(\bar{x}\) than for \(X\). A cutoff below \(\mu\) will usually have a larger probability of being exceeded by \(\bar{x}\).

A Worked Comparison

The main task is to identify what the random variable represents before choosing a distribution. Then use the relevant spread and find the probability to the right of the cutoff. The following example uses one population and one cutoff to compare both questions directly.

1
State.
Define whether the event concerns one observation \(X\) or the sample mean \(\bar{x}\), and identify the cutoff \(c\).
2
Plan.
Identify the appropriate distribution and standard deviation. Check the sampling assumptions and whether the Normal model is exact or approximate.
3
Do.
Standardize the cutoff or use normalcdf with the correct mean and standard deviation for each random variable.
4
Conclude.
Interpret each probability in context, making clear whether it describes individual observations or sample means.

Worked Example: Individual Bottle Fill or Sample Mean?

Suppose fill volumes from a bottling process follow a Normal population with mean \(\mu=500\) milliliters and standard deviation \(\sigma=12\) milliliters. Individual fills are independent. Compare the probability that one fill exceeds 506 milliliters with the probability that the mean of 36 fills exceeds 506 milliliters.

State. Let \(X\) be the fill volume, in milliliters, of one bottle, and let \(\bar{x}\) be the mean fill volume for a random sample of 36 bottles. We want to compare \(P(X>506)\) and \(P(\bar{x}>506)\).

Plan. The population is Normal, so \(X\) is exactly Normal with mean 500 and standard deviation 12. The observations are independent, so \(\bar{x}\) is exactly Normal with mean 500 and standard deviation \(\sigma/\sqrt{n}\). These independent observations are drawn from a production process; if instead the sample were selected without replacement from a finite batch, we would check the 10% condition \(n\leq0.10N\).

Do: one bottle. Use the population standard deviation for an individual fill. Its standardized cutoff is:

$$ z=\frac{506-500}{12}=\frac{6}{12}=0.5 $$

Using normalcdf with lower bound 506, a large upper bound, mean 500, and standard deviation 12 gives:

$$ P(X>506)=\operatorname{normalcdf}(506,\infty,500,12) \approx 0.3085 $$

As a check, the standard Normal upper-tail probability above \(z=0.5\) is \(1-0.6915=0.3085\), rounded to four decimal places.

Do: mean of 36 bottles. First find the standard deviation of the sample mean:

$$ \sigma_{\bar{x}}=\frac{12}{\sqrt{36}}=\frac{12}{6}=2\text{ milliliters} $$

The variance check gives \(12^2/36=144/36=4\) square milliliters, and \(\sqrt{4}=2\) milliliters. Now standardize the same cutoff using the sampling-distribution standard deviation:

$$ z=\frac{506-500}{2}=\frac{6}{2}=3 $$

Using normalcdf with mean 500 and standard deviation 2 gives:

$$ P(\bar{x}>506)=\operatorname{normalcdf}(506,\infty,500,2) \approx 0.0013 $$

The standard Normal upper-tail probability above \(z=3\) is about 0.00135, which rounds to 0.0013 to four decimal places.

Conclude. Under the stated model, about 30.85% of individual bottle fills exceed 506 milliliters, while only about 0.13% of sample means from samples of 36 bottles exceed 506 milliliters. The cutoff is three standard deviations above the center of the sample-mean distribution but only half a standard deviation above the center of the individual-value distribution.

How the Cutoff’s Location Matters

The bottling example used a cutoff above the population mean. Because \(\bar{x}\) has a smaller standard deviation than \(X\), sample means are less likely to reach a high cutoff. If a cutoff is below the population mean, the comparison goes in the opposite direction: sample means are more likely than individual values to be above that low cutoff.

For a symmetric Normal population, a cutoff equal to \(\mu\) gives an upper-tail probability of 0.5 for both \(X\) and \(\bar{x}\). This is because both Normal distributions are centered at \(\mu\). The probabilities differ when the cutoff is away from the center, and how much they differ depends on the sample size as well as the cutoff.

Cutoff positionTypical comparison for \(P(X>c)\) and \(P(\bar{x}>c)\)
\(c>\mu\)\(P(\bar{x}>c)\) is smaller because the sample-mean distribution is less spread out.
\(c=\mu\)For symmetric Normal distributions, both probabilities are 0.5.
\(c<\mu\)\(P(\bar{x}>c)\) is larger because most sample means cluster near \(\mu\), above the low cutoff.

Worked Example: A Less Extreme Cutoff for Bottle Fills

Use the same bottling population, with \(\mu=500\) milliliters and \(\sigma=12\) milliliters, but change the cutoff to 502 milliliters. Compare the chance that one fill and the mean of 36 fills exceed this cutoff.

State. \(X\) represents one fill, and \(\bar{x}\) represents the mean fill for a sample of 36 independent bottles. We seek \(P(X>502)\) and \(P(\bar{x}>502)\).

Plan. The Normal population gives an exact Normal model for \(X\) and, with independent observations, an exact Normal model for \(\bar{x}\). Use standard deviations 12 milliliters and 2 milliliters, respectively. The sample-mean standard deviation is \(12/\sqrt{36}=2\) milliliters.

Do: one bottle. The individual-value \(z\)-score and upper-tail probability are:

$$ z_X=\frac{502-500}{12}=\frac{2}{12}\approx0.1667 $$

Therefore, \(\operatorname{normalcdf}(502,\infty,500,12)\approx0.4338\). A check using the standard Normal distribution gives \(1-\Phi(0.1667)\approx1-0.5662=0.4338\), rounded to four decimal places.

Do: mean of 36 bottles. The sample-mean \(z\)-score and upper-tail probability are:

$$ z_{\bar{x}}=\frac{502-500}{2}=1 $$

Thus, \(\operatorname{normalcdf}(502,\infty,500,2)\approx0.1587\). The check is \(1-\Phi(1)=1-0.8413=0.1587\), rounded to four decimal places.

Conclude. About 43.38% of individual fills exceed 502 milliliters, compared with about 15.87% of means from samples of 36 fills. The cutoff is above the center in both distributions, but it is farther above the center when measured in standard deviations for \(\bar{x}\).

Worked Example: A Cutoff Below the Population Mean

Suppose the time a device takes to complete a startup sequence follows a Normal population with mean \(\mu=40\) seconds and standard deviation \(\sigma=8\) seconds. Startup times are independent. Compare the chance that one startup takes more than 38 seconds with the chance that the mean startup time for 16 sequences exceeds 38 seconds.

State. Let \(X\) be the startup time, in seconds, for one sequence, and let \(\bar{x}\) be the mean startup time for 16 independent sequences. The requested probabilities are \(P(X>38)\) and \(P(\bar{x}>38)\).

Plan. The population is Normal, so the model for \(X\) is exact. Independence makes the sampling distribution of \(\bar{x}\) exactly Normal as well. Its standard deviation is \(\sigma/\sqrt{n}\). If these 16 sequences were sampled without replacement from a finite collection, we would also check that \(16\leq0.10N\).

Do: one startup. The cutoff is below the mean. For one observation:

$$ z_X=\frac{38-40}{8}=\frac{-2}{8}=-0.25 $$

The probability above this cutoff is \(\operatorname{normalcdf}(38,\infty,40,8)\approx0.5987\). Equivalently, \(1-\Phi(-0.25)=1-0.4013=0.5987\), rounded to four decimal places.

Do: mean of 16 startups. First calculate and verify the standard deviation of the sample mean:

$$ \sigma_{\bar{x}}=\frac{8}{\sqrt{16}}=\frac{8}{4}=2\text{ seconds} $$

The variance check is \(8^2/16=64/16=4\) square seconds, whose square root is 2 seconds. The standardized cutoff and upper-tail probability are:

$$ z_{\bar{x}}=\frac{38-40}{2}=-1 $$

Thus, \(\operatorname{normalcdf}(38,\infty,40,2)\approx0.8413\), or \(1-\Phi(-1)=1-0.1587=0.8413\), rounded to four decimal places.

Conclude. About 59.87% of individual startup times exceed 38 seconds, while about 84.13% of means from samples of 16 sequences exceed 38 seconds. Since 38 seconds is below the population mean, the tighter sample-mean distribution places more of its area above this cutoff.

What the CLT Does—and Does Not—Tell You

The examples used a Normal population, so both the individual-value and sample-mean models were exactly Normal. When a population is not Normal, take care: the Central Limit Theorem concerns the sampling distribution of \(\bar{x}\), not the distribution of individual observations \(X\). As discussed in “Central Limit Theorem Explained” and “Is \(n=30\) Large Enough for the CLT,” the population shape and sample size determine whether a Normal approximation for sample means is reasonable.

A large sample does not make individual observations Normal. To calculate \(P(X>c)\), use a Normal model only if the population itself is Normal or if a suitable model for individual values is otherwise justified. A CLT-based approximation might support calculating \(P(\bar{x}>c)\), but it does not support calculating \(P(X>c)\).

For random samples taken without replacement, check the 10% condition \(n\leq0.10N\) before treating observations as independent. For independent observations by design, state that independence is assumed. These checks matter for the sample-mean calculation because its sampling distribution depends on the sampling process.

Common Mistakes and AP Exam Tip

  • Using \(\sigma/\sqrt{n}\) for an individual value: \(X\) describes one observation, so use the population standard deviation \(\sigma\). The standard deviation \(\sigma/\sqrt{n}\) belongs to \(\bar{x}\).
  • Using \(\sigma\) for a sample mean: Averages vary less than individual observations under the stated independence conditions. For \(\bar{x}\), use \(\sigma/\sqrt{n}\), not \(\sigma\).
  • Mixing up the events: \(P(X>c)\) describes an individual measurement exceeding \(c\). \(P(\bar{x}>c)\) describes a sample average from samples of a specified size exceeding \(c\). Name the random variable in the conclusion.
  • Using the CLT for individual values: The CLT helps describe the sampling distribution of a mean; it does not turn the population distribution of \(X\) into a Normal distribution.
  • Assuming the sample-mean probability is always smaller: That is generally true for a cutoff above \(\mu\), but not for a cutoff below \(\mu\). Compare the cutoff’s location with the center before predicting the result.
  • Reporting a calculator value without explaining its inputs: Show the mean and standard deviation used for each distribution, and include the context and units in the interpretation.

For full-credit communication, define \(X\) and \(\bar{x}\), state the relevant population or sampling-distribution parameters, justify the model, and interpret each probability with its subject and context. Make clear that the same cutoff does not make the two probability questions equivalent.

Key takeaway: \(P(X>c)\) uses the distribution of individual observations and spread \(\sigma\); \(P(\bar{x}>c)\) uses the sampling distribution of sample means and spread \(\sigma/\sqrt{n}\). Compare the cutoff with \(\mu\) to understand why the probabilities differ.

Check Your Understanding

For each question, identify the random variable, select the appropriate standard deviation, and explain what the probability describes.

  1. A Normal population has mean 72 centimeters and standard deviation 10 centimeters. Compare \(P(X>77)\) with \(P(\bar{x}>77)\) for independent samples of size 25. Which probability is smaller, and why?
  2. A Normal population has mean 18 minutes and standard deviation 6 minutes. For independent samples of size 9, find \(P(X>15)\) and \(P(\bar{x}>15)\). Interpret both probabilities.
  3. Individual observations come from a strongly skewed population. A student says the CLT makes \(X\) approximately Normal for a sample of size 100. Explain the error.
  4. For a symmetric Normal population, what are \(P(X>\mu)\) and \(P(\bar{x}>\mu)\)? Explain why they are equal.
  5. A student calculates \(P(\bar{x}>c)\) using \(\sigma\) instead of \(\sigma/\sqrt{n}\). Describe how this confuses the individual-value distribution with the sampling distribution.