Tutorials › AP Statistics › Normal Model for Sample Proportion Probabilities

Sampling distributions for proportions · Tutorial 410 of 1000

Normal Model for Sample Proportion Probabilities

Build a Normal model for \(\hat{p}\), then identify the mean, standard deviation, and curve region that match a stated probability.

Intermediate 9 min read

What You'll Learn

  • Write the Normal model for a sample proportion using the population proportion and the standard deviation of \(\hat{p}\).
  • Check the randomization, 10% condition, and Large Counts condition before using the model.
  • Translate probability statements about \(\hat{p}\) into left-tail, right-tail, or between-values regions.
  • Label the mean, cutoff values, and shaded region on a Normal curve sketch.
  • Distinguish the Normal model for \(\hat{p}\) from the distribution of individual outcomes in a sample.

From a Sampling Distribution to a Normal Model

In Shape of the Sampling Distribution When \(p\) Is Near 0 or 1, you saw why the sampling distribution of \(\hat{p}\) may be skewed when expected counts are small. When the conditions support a Normal approximation, we can describe sample-proportion probabilities as areas under a Normal curve. This tutorial focuses on setting up that model and showing which area represents a stated event.

The model is for \(\hat{p}\), the sample proportion across repeated random samples of the same size—not for the outcome of one individual. As established in Mean of the Sampling Distribution of p-hat and Standard Deviation of p-hat Formula, the sampling distribution has mean \(p\) and standard deviation \(\sqrt{p(1-p)/n}\), when the sampling conditions are appropriate.

Formula: When a Normal approximation is appropriate, model the sample proportion as:
$$ \hat{p}\ \approx\ N\left(p,\sqrt{\frac{p(1-p)}{n}}\right) $$
In this course, \(N(\text{mean},\text{standard deviation})\) lists the mean first and the standard deviation second. Thus, the center is \(p\), and the spread is \(\sigma_{\hat{p}}=\sqrt{p(1-p)/n}\).

The model is an approximation: the actual sample proportion can take only values in steps of \(1/n\), while a Normal curve is continuous. The curve is useful when the conditions support it, but it does not make the actual sampling distribution exactly Normal. It is also possible for a Normal curve to extend slightly beyond 0 or 1, even though a sample proportion cannot; the approximation is useful only when it reasonably represents the values that can occur.

Check Conditions Before Drawing the Curve

Before sketching or using a Normal model, make the sampling assumptions visible. As in the earlier tutorials on random sampling, the 10% condition, and checking normality of \(\hat{p}\) with \(np\) and \(n(1-p)\), check that the process is random or otherwise justifies treating observations as random, that observations are independent, and that both expected counts are large enough.

Conditions: For a Normal model of \(\hat{p}\), check that:
  • The sample is random, or the sampling process otherwise supports treating observations as random.
  • If sampling without replacement from a finite population of size \(N\), the 10% condition holds: \(n\leq0.10N\).
  • The Large Counts condition holds: \(np\geq10\) and \(n(1-p)\geq10\).
When sampling with replacement or under an appropriate independent-trials model, the 10% condition is not needed. State the relevant independence justification for the setting.

Once the model is supported, translate the event into a region. A statement that \(\hat{p}\) is below a cutoff corresponds to the area to the left; above a cutoff corresponds to the area to the right; and between two cutoffs corresponds to the area between them. This is the same area logic used for other Normal random variables in Finding Areas Above and Below a Value and Probability Between Two Values on a Normal Curve. The important extra step here is using the correct mean and standard deviation for \(\hat{p}\).

1
Identify the sample proportion and event.
Write what \(\hat{p}\) measures and translate the probability statement into an inequality or interval.
2
Check the model conditions.
Give evidence for randomization or randomness, check the 10% condition when relevant, and calculate both expected counts.
3
Write the Normal model.
Use mean \(p\) and standard deviation \(\sqrt{p(1-p)/n}\), keeping the parameters in that order.
4
Sketch and label the requested region.
Draw a bell-shaped curve, mark its center \(p\), place the event cutoff or cutoffs on the horizontal axis, and shade the matching area.

Upper-Tail Probabilities

For an upper-tail event such as \(P(\hat{p}\geq c)\), the boundary \(c\) determines where the shaded region begins. Put \(p\) at the center of the curve and \(c\) at its location on the horizontal axis. Shade to the right of \(c\). Do not shade from the center unless the cutoff actually equals the mean.

Worked Example: At Least 48% Support in a Sample

A community group plans a random sample of 200 adults from a town of 10,000 adults. Suppose 42% of all adults in the town support a proposed park. Find the Normal model for \(\hat{p}\), the sample proportion who support the proposal, and set up the region for the probability that at least 48% of the sample supports it.

State. Let \(\hat{p}\) be the proportion of sampled adults who support the proposal. The event is \(P(\hat{p}\geq0.48)\), and the population proportion is \(p=0.42\).

Plan and check conditions. The sample is described as random. Because it is drawn without replacement from a finite population, check the 10% condition: \(200\leq0.10(10{,}000)=1{,}000\), so it is met. The expected success and failure counts are \(np=200(0.42)=84\) and \(n(1-p)=200(0.58)=116\). Both are at least 10, so the Large Counts condition is met. The Normal model is supported.

Do: calculate the model. The mean is \(p=0.42\). The standard deviation is:

$$ \sigma_{\hat{p}} =\sqrt{\frac{0.42(1-0.42)}{200}} =\sqrt{\frac{0.42(0.58)}{200}} =\sqrt{0.001218} \approx0.03490 $$

Therefore, \(\hat{p}\) is modeled approximately by \(N(0.42,0.03490)\). For the sketch, draw a bell-shaped curve centered at \(0.42\). Mark \(0.48\) to the right of the center, and shade the curve to the right of \(0.48\). Label the shaded area \(P(\hat{p}\geq0.48)\). This identifies the requested probability as a right-tail area; evaluating the area is a separate calculation.

Conclude. Under the stated sampling assumptions, the Normal model with mean \(0.42\) and standard deviation about \(0.03490\) represents the sampling distribution of the support proportion. The region to the right of \(0.48\) represents the chance that a random sample of 200 adults has a support proportion of at least 48%.

Probabilities Between Two Cutoffs

A probability statement that places \(\hat{p}\) between a lower and an upper value corresponds to the area between those values. Mark both cutoffs on the horizontal axis and shade only the section of the curve between them. Check their positions relative to the mean: an interval centered at \(p\) should have the same distance to either side of the center.

Worked Example: A Sample Proportion Within a Range

A college surveys a random sample of 160 students from a student population of 5,000. Suppose 70% of the population uses a particular study app. Let \(\hat{p}\) be the proportion of sampled students who use the app. Set up a Normal model and identify the region for \(P(0.64\leq\hat{p}\leq0.76)\).

State. The parameter is \(p=0.70\), and the event asks for a sample proportion from 0.64 through 0.76, inclusive.

Plan and check conditions. The survey uses a random sample. Since the sample is taken without replacement, check the 10% condition: \(160\leq0.10(5{,}000)=500\), so it is met. The expected number of app users is \(np=160(0.70)=112\). The expected number of nonusers is \(n(1-p)=160(0.30)=48\). Both counts are at least 10, meeting the Large Counts condition. A Normal model is appropriate for the sampling distribution.

Do: calculate the model and identify the area. The mean is \(0.70\). The standard deviation is:

$$ \sigma_{\hat{p}} =\sqrt{\frac{0.70(1-0.70)}{160}} =\sqrt{\frac{0.70(0.30)}{160}} =\sqrt{0.0013125} \approx0.03623 $$

Thus, \(\hat{p}\) is modeled approximately by \(N(0.70,0.03623)\). On the sketch, center the curve at \(0.70\). Mark \(0.64\) to the left and \(0.76\) to the right; each is \(0.06\) from the center. Shade the area between those two marks and label it \(P(0.64\leq\hat{p}\leq0.76)\). Do not shade the tails outside the interval.

Conclude. The shaded area represents the probability that a random sample of 160 students has an app-use proportion between 64% and 76%, under the Normal approximation and the stated population and sampling assumptions.

Lower-Tail Probabilities and Cutoffs Below the Mean

For a lower-tail event, shade from the far left up to the cutoff. When the cutoff is below the mean, the shaded region will be less than half the curve; when it is above the mean, it will be more than half. This is a useful visual check that the sketch matches the inequality.

Worked Example: A Sample Proportion at Least 15%

A conservation team takes a random sample of 500 plants from a very large field. Suppose 12% of the plants in the field have a particular leaf condition. Let \(\hat{p}\) be the proportion in the sample with that condition. Set up the model and sketch the region for \(P(\hat{p}\geq0.15)\).

State. Here \(p=0.12\), \(n=500\), and the event is that the sample proportion is at least 0.15.

Plan and check conditions. The sample is described as random. The field is very large relative to the sample, so sampling without replacement changes the population by no more than 10%; the 10% condition is reasonable. The expected number of plants with the condition is \(np=500(0.12)=60\), and the expected number without it is \(n(1-p)=500(0.88)=440\). Both are at least 10, so the Large Counts condition is met. The Normal approximation is supported.

Do: calculate the model. The mean is \(0.12\), and the standard deviation is:

$$ \sigma_{\hat{p}} =\sqrt{\frac{0.12(1-0.12)}{500}} =\sqrt{\frac{0.12(0.88)}{500}} =\sqrt{0.0002112} \approx0.01453 $$

The model is approximately \(N(0.12,0.01453)\). Draw a bell-shaped curve centered at \(0.12\), and mark \(0.15\) to the right of the center. Shade from \(0.15\) toward the right end of the curve, leaving the area to the left unshaded. Label the region \(P(\hat{p}\geq0.15)\). Since the cutoff is above the mean, the right-tail area should be less than one-half.

Conclude. The shaded right tail represents the probability that a random sample of 500 plants has a condition proportion of at least 15%, assuming the stated model and sampling process.

Reading the Sketch Before Calculating

A sketch is not decoration: it helps verify that the probability statement, inequality, and calculator inputs will agree. The curve’s center must be \(p\), not the observed sample proportion from a particular sample. The curve’s horizontal scale is in units of \(\hat{p}\), so the cutoffs must be proportions such as \(0.15\), not counts such as 75 plants. If a problem gives a count, first divide by \(n\) to express the event as a sample proportion.

The event’s wording determines the shaded side. “At most” or “no more than” points left; “at least” or “no fewer than” points right; “between” points to the interior. A strict or inclusive inequality does not change the area under a continuous Normal model at a single boundary, since a single point has area zero. But keep the inequality as stated when writing the event and label.

The standard deviation also needs careful attention. It describes the variability of sample proportions from sample to sample, not the variability of individual 0-or-1 outcomes. Use \(\sqrt{p(1-p)/n}\), and place that result second in \(N(p,\sigma_{\hat{p}})\). Do not substitute \(\sqrt{p(1-p)}\) without dividing by \(n\) inside the square root.

Common Mistakes and AP Exam Communication

  • Using the wrong center. The Normal model for \(\hat{p}\) is centered at the population proportion \(p\), not at a sample proportion from observed data.
  • Putting the parameters in the wrong order. In the course convention \(N(\text{mean},\text{standard deviation})\), write \(N(p,\sigma_{\hat{p}})\), not the reverse.
  • Using the standard deviation of individual outcomes. For \(\hat{p}\), use \(\sqrt{p(1-p)/n}\). The spread shrinks as \(n\) increases.
  • Checking only one expected count. Show both \(np\) and \(n(1-p)\); both must be at least 10 for the Large Counts condition.
  • Shading the wrong region. A “at least” statement is a right tail; a “at most” statement is a left tail; a “between” statement is the middle interval.
  • Forgetting the event’s units. A cutoff for \(\hat{p}\) is a proportion, such as \(0.48\). If the statement gives a number of individuals, convert the count to a proportion before locating it on the \(\hat{p}\) axis.
  • Calling the approximation exact. Say the sampling distribution is modeled approximately by a Normal distribution when the conditions support the approximation.
AP Exam Tip: Make the reasoning easy to follow: define \(\hat{p}\), check the relevant sampling and Large Counts conditions, show the standard deviation calculation, write \(N(p,\sigma_{\hat{p}})\), and describe the shaded region using the event’s inequality. A labeled sketch should include the center, every cutoff, and exactly the requested area.

Key Takeaway

To represent probabilities about a sample proportion, first check whether a Normal approximation is supported. Then use \(p\) as the mean and \(\sqrt{p(1-p)/n}\) as the standard deviation. Translate the event into a left tail, right tail, or interval, and make the sketch match that event exactly.

Key takeaway: For an appropriate Normal approximation, \(\hat{p}\approx N\left(p,\sqrt{p(1-p)/n}\right)\). Label the center \(p\), mark the event’s cutoff or cutoffs, and shade the area specified by the probability statement.

Check Your Understanding

For each question, write or describe the model and identify the correct region on a labeled Normal curve.

  1. A random sample of 250 people is drawn from a population with \(p=0.36\). The population is large. Check the Large Counts condition, calculate the standard deviation of \(\hat{p}\), and set up the model.
  2. For the setting in question 1, describe the region for \(P(\hat{p}\leq0.32)\). Where is the cutoff relative to the center?
  3. A random sample of 120 customers is drawn without replacement from a group of 900. Suppose \(p=0.65\). Check the 10% condition and both expected counts. Is the Normal model supported?
  4. For the setting in question 3, describe the sketch for \(P(0.60\leq\hat{p}\leq0.70)\), including the center and both cutoffs.
  5. A random sample of 400 seeds comes from a population with \(p=0.08\). State the expected success and failure counts and say whether the Large Counts condition supports a Normal model.