Tutorials › AP Statistics › Normal Models for Sample Means and Proportions

Normal distributions · Tutorial 376 of 1000

Normal Models for Sample Means and Proportions

Learn to model the sampling distributions of sample means and proportions, check their conditions, and use standard errors with normal areas to find probabilities.

Intermediate 9 min read

What You'll Learn

  • Distinguish the distribution of individual observations from the sampling distribution of a sample mean.
  • Calculate the standard error of a sample mean and interpret it in context.
  • Check when a normal model is appropriate for a sample mean.
  • Find probabilities for sample means using normalcdf and the sampling distribution’s standard error.
  • Calculate the standard error of a sample proportion and check the Large Counts condition.
  • Use a normal model to estimate probabilities for sample proportions.

From Individual Values to Sample Statistics

A measurement from one individual can vary considerably, so a probability about one individual may differ from a probability about the average of a sample. The average, \(\bar{x}\), changes from sample to sample too, but its sampling distribution generally has less spread than the distribution of individual measurements. In this tutorial, we use a normal model for that sampling distribution to find probabilities about sample means. We also apply the same idea to sample proportions, \(\hat{p}\).

The sampling distribution of a statistic describes the values that statistic would take across all possible random samples of the same size from a population. Its standard deviation is called the statistic’s standard error. A standard error describes the typical sample-to-sample variation in a statistic; it is not the standard deviation of individual observations.

Formula: For random samples of size \(n\) with independent observations, or when the 10% condition justifies treating observations sampled without replacement as approximately independent, the sampling distribution of the sample mean \(\bar{x}\) has mean \(\mu_{\bar{x}}=\mu\) and standard error \(\sigma_{\bar{x}}=\frac{\sigma}{\sqrt{n}}\). For a simple random sample without replacement from a finite population of size \(N\), when the sampling fraction is not negligible, the standard error is \(\frac{\sigma}{\sqrt{n}}\sqrt{\frac{N-n}{N-1}}\), where \(\sigma\) is the population standard deviation. The sampling distribution of the sample proportion \(\hat{p}\) has mean \(\mu_{\hat{p}}=p\) and standard error \(\sigma_{\hat{p}}=\sqrt{\frac{p(1-p)}{n}}\), when the conditions for using these models are met.

The sample mean is centered at the population mean, and the sample proportion is centered at the population proportion. In both cases, increasing the sample size reduces the standard error. For a sample mean, the decrease follows \(1/\sqrt{n}\); for a sample proportion, it also follows \(1/\sqrt{n}\) when the population proportion stays fixed.

When a Normal Model Is Appropriate

A normal model for \(\bar{x}\) is appropriate when the population itself is normally distributed, or when the sample size is large enough for the Central Limit Theorem to make the sampling distribution of \(\bar{x}\) approximately normal. How large is “large enough” depends on the population’s shape: a strongly skewed population or extreme outliers can require a larger sample. Also, if observations are sampled without replacement, check the 10% condition to justify treating them as independent.

Conditions for a normal model of \(\bar{x}\):
  • The data come from a random sample or a process that justifies treating the observations as random.
  • If sampling without replacement, the population is at least 10 times the sample size: \(N\geq10n\). This is the 10% condition.
  • The population is normal, or the sample is large enough for the sampling distribution of \(\bar{x}\) to be approximately normal.

For \(\hat{p}\), the population outcomes must be binary: each individual either has the characteristic defined as a success or does not. When the sample is random and the independence condition is reasonable, a normal model for the sampling distribution is appropriate if the Large Counts condition holds.

Conditions for a normal model of \(\hat{p}\):
  • The data come from a random sample or a process that justifies treating the observations as random.
  • If sampling without replacement, check the 10% condition, \(N\geq10n\).
  • The Large Counts condition holds: \(np\geq10\) and \(n(1-p)\geq10\).

Once conditions support the normal model, use the relevant mean and standard error as its parameters. For a sample mean, those are \(\mu\) and \(\sigma/\sqrt{n}\). For a sample proportion, they are \(p\) and \(\sqrt{p(1-p)/n}\). Enter the event’s bounds in the original units of the statistic. For example, a bound for \(\bar{x}\) is in the same units as the individual measurements, while a bound for \(\hat{p}\) is a proportion.

Finding Probabilities for Sample Means

A useful way to organize a sample-mean probability is to first identify the sampling distribution and calculate its standard error. Then sketch or state the event and find its area under the corresponding normal curve. As in Using normalcdf to Find a Normal Area, the calculator inputs use the lower bound, upper bound, mean, and standard deviation—in this case, the standard error is the sampling distribution’s standard deviation.

1
Define the statistic and check conditions.
State what \(\bar{x}\) measures and its units. Check randomness, the 10% condition when needed, and normality or a sufficiently large sample.
2
Find the sampling distribution’s parameters.
Use mean \(\mu\) and standard error \(\sigma/\sqrt{n}\).
3
Translate the event and find its area.
Use the given bounds for \(\bar{x}\) in \(\operatorname{normalcdf}(\text{lower},\text{upper},\mu,\sigma/\sqrt{n})\).
4
Interpret in context.
Describe the probability as the chance that a random sample of the stated size has a sample mean in the requested range.

Worked Example: Probability About an Average Battery Life

Worked Example: Probability About an Average Battery Life

In a fictional quality-control setting, battery life has a normal population distribution with mean 8.4 hours and standard deviation 1.2 hours. A random sample of 36 batteries is selected. Find the probability that the sample’s mean battery life is between 8.3 and 8.6 hours.

State. Let \(\bar{X}\) be the mean battery life, in hours, for a random sample of 36 batteries. We want \(P(8.3\leq\bar{X}\leq8.6)\).

Plan. The population is stated to be normal, and the sample is random. Assume the population contains at least 360 batteries, so the 10% condition holds if sampling is without replacement. The sampling distribution of \(\bar{X}\) is normal, with mean 8.4 hours and standard error \(1.2/\sqrt{36}\) hours.

Do. Calculate the standard error, then find the area between the two sample-mean bounds:

$$ \sigma_{\bar{x}}=\frac{\sigma}{\sqrt{n}} =\frac{1.2}{\sqrt{36}} =\frac{1.2}{6} =0.2\text{ hours} $$
$$ P(8.3\leq\bar{X}\leq8.6) =\operatorname{normalcdf}(8.3,8.6,8.4,0.2) \approx0.5328 $$

As a check using standardized boundaries, \((8.3-8.4)/0.2=-0.5\) and \((8.6-8.4)/0.2=1.0\). The area between \(z=-0.5\) and \(z=1.0\) is approximately \(0.8413-0.3085=0.5328\), rounded to four decimal places.

Conclude. The probability that a random sample of 36 batteries has a mean life between 8.3 and 8.6 hours is approximately 0.5328. This is a probability about sample averages, not the proportion of individual batteries in that interval.

Worked Example: A Sample Mean from a Skewed Population

Worked Example: A Sample Mean from a Skewed Population

In a fictional community, individual daily water use is right-skewed, with population mean 42 liters and standard deviation 16 liters. Random samples of 64 households are taken. Assuming this sample size is large enough for the sampling distribution of the mean to be approximately normal, estimate the probability that a sample’s mean daily use exceeds 45 liters.

State. Let \(\bar{X}\) be the mean daily water use, in liters, for a random sample of 64 households. The event is \(P(\bar{X}>45)\).

Plan. The population is skewed, so we rely on the stated assumption that \(n=64\) is sufficiently large for the Central Limit Theorem to give an approximately normal sampling distribution. The sample is random. Assume there are at least 640 households if sampling without replacement, satisfying the 10% condition. The sampling distribution has mean 42 liters and standard error \(16/\sqrt{64}\) liters.

Do. Find the standard error and the right-tail area:

$$ \sigma_{\bar{x}}=\frac{16}{\sqrt{64}} =\frac{16}{8} =2\text{ liters} $$
$$ P(\bar{X}>45) =\operatorname{normalcdf}(45,1E99,42,2) \approx0.0668 $$

The standardized boundary is \((45-42)/2=1.5\). The right-tail area above \(z=1.5\) is approximately 0.0668, rounded to four decimal places, agreeing with the calculator result.

Conclude. Under the approximate normal model, the probability that a random sample of 64 households has mean daily water use greater than 45 liters is about 0.0668. The skewness of individual household use does not disappear from the population, but the sampling distribution of the mean is modeled as approximately normal under the stated large-sample assumption.

Finding Probabilities for Sample Proportions

A sample proportion is the fraction of sampled individuals who have the characteristic defined as a success. For instance, if 60 out of 100 sampled households use a particular service, then \(\hat{p}=60/100=0.60\). The population proportion \(p\) is the long-run fraction of the population with that characteristic. When the conditions are met, use \(p\) as the mean of the sampling distribution and \(\sqrt{p(1-p)/n}\) as its standard error.

As with sample means, the probability describes how often sample statistics would fall in a specified range over repeated random sampling under the model. It is not a claim that the population proportion itself varies from sample to sample. Here, \(p\) is treated as a known population value for calculating the probability.

Worked Example: A Sample Proportion Above a Cutoff

Worked Example: A Sample Proportion Above a Cutoff

In a fictional town, 50% of households have a home garden. A random sample of 100 households is selected. Find the approximate probability that at least 60% of the sampled households have a home garden.

State. Let \(\hat{p}\) be the proportion of sampled households that have a home garden. We want \(P(\hat{p}\geq0.60)\), with population proportion \(p=0.50\).

Plan. The sample is random. Assume the town has at least 1,000 households, so the 10% condition holds for sampling without replacement. The Large Counts condition is satisfied because \(np=100(0.50)=50\geq10\) and \(n(1-p)=100(0.50)=50\geq10\). Therefore, a normal model for \(\hat{p}\) is appropriate.

Do. Calculate the mean and standard error of the sampling distribution, then find the right-tail area:

$$ \mu_{\hat{p}}=p=0.50 \qquad \sigma_{\hat{p}}=\sqrt{\frac{p(1-p)}{n}} =\sqrt{\frac{(0.50)(0.50)}{100}} =\sqrt{0.0025} =0.05 $$
$$ P(\hat{p}\geq0.60) \approx\operatorname{normalcdf}(0.60,1E99,0.50,0.05) \approx0.0228 $$

The standardized boundary is \((0.60-0.50)/0.05=2\). The right-tail area above \(z=2\) is approximately 0.0228, rounded to four decimal places. The result is an approximation because the sampling distribution is modeled with a continuous normal curve.

Conclude. If the population proportion of households with a home garden is 0.50, the approximate probability that at least 60% of a random sample of 100 households have a home garden is 0.0228.

Common Mistakes and AP Exam Tips

  • Using the individual standard deviation for a sample mean. For \(\bar{x}\), use the standard error \(\sigma/\sqrt{n}\), not \(\sigma\), as the normal model’s standard deviation.
  • Forgetting to check conditions. A normal area calculation does not establish that a normal model is appropriate. State why the sampling is random, check the 10% condition when applicable, and check normality or the Large Counts condition as relevant.
  • Mixing up \(p\) and \(\hat{p}\). The population proportion \(p\) is the center of the sampling distribution; \(\hat{p}\) is the random sample statistic whose probability is being calculated.
  • Using the wrong standard error for a proportion. Use \(\sqrt{p(1-p)/n}\), not \(p(1-p)/n\). The latter is the variance, not the standard deviation.
  • Entering the sample size as the standard deviation. In normalcdf, the fourth input is the standard error of the statistic, not \(n\) or the population size.
  • Interpreting the probability as a statement about individuals. A probability involving \(\bar{x}\) is about sample averages; one involving \(\hat{p}\) is about sample proportions.
  • Reporting an approximation as exact. When the sampling distribution is only approximately normal, say the result is approximate and report the probability rounded appropriately.

A full-credit response defines the statistic in context, checks the conditions explicitly, identifies the sampling distribution’s mean and standard error, shows the normal-area calculation, and answers the question in context. Keep units with a sample mean and its standard error. A proportion and its standard error are expressed as proportions, not in the units of individual observations.

AP Exam Tip: The sample size changes the standard error, not the center: \(\bar{x}\) remains centered at \(\mu\), and \(\hat{p}\) remains centered at \(p\). Before using normalcdf, write down the statistic’s mean and standard error separately.

Key Takeaway

Normal areas can describe how a sample mean or sample proportion varies over repeated random samples. First check that the sampling distribution is approximately normal; then use its own standard error as the standard deviation in the normal probability calculation.

Key takeaway: For a sample mean, use mean \(\mu\) and standard error \(\sigma/\sqrt{n}\). For a sample proportion, use mean \(p\) and standard error \(\sqrt{p(1-p)/n}\). Check the appropriate conditions, calculate the normal area, and interpret it as a probability about sample statistics.

Check Your Understanding

For each question, identify the statistic’s sampling distribution, check the relevant conditions, and show how you would find the probability.

  1. A normally distributed measurement has mean 72 centimeters and standard deviation 10 centimeters. For a random sample of 25 measurements, what are the mean and standard error of the sample-mean sampling distribution?
  2. A population has mean 18 minutes and standard deviation 6 minutes. A random sample of 36 observations is taken. Assuming the sample-mean distribution is approximately normal and the 10% condition holds, write the normalcdf expression for the probability that \(\bar{x}\) is below 17 minutes.
  3. In a population, the proportion with a certain feature is \(p=0.20\). For a random sample of 80, check the Large Counts condition and calculate the standard error of \(\hat{p}\).
  4. A random sample of 100 people is taken from a large population where \(p=0.40\). Assuming the 10% condition holds, write the normalcdf expression for the approximate probability that \(\hat{p}\) is between 0.35 and 0.45.
  5. Explain why the standard deviation of individual observations is not used directly as the standard deviation of the sampling distribution of \(\bar{x}\).