Tutorials › AP Statistics › Shape of the Sampling Distribution When p Is Near 0 or 1

Sampling distributions for proportions · Tutorial 409 of 1000

Shape of the Sampling Distribution When p Is Near 0 or 1

Compare the shape of the sampling distribution of \(\hat{p}\) for \(p=0.05\) at small and large sample sizes, and see why larger samples improve the Normal approximation.

Intermediate 10 min read

What You'll Learn

  • Explain why a small expected number of successes can make the sampling distribution of \(\hat{p}\) right-skewed.
  • Connect possible sample proportions to the binomial count of successes.
  • Compare the shape and spread for \(p=0.05\) at different sample sizes.
  • Describe how increasing \(n\) makes the distribution more nearly symmetric without making it exactly Normal.
  • Explain how skewness changes when the population proportion is near 1 rather than near 0.
  • Distinguish the Large Counts check from a guarantee about distribution shape.

Why a Small Proportion Can Produce a Skewed Sampling Distribution

In Checking Normality of \(\hat{p}\) with \(np\) and \(n(1-p)\), you learned to use the Large Counts condition to check whether a Normal model is supported. This tutorial looks at the shape behind that check. When the population proportion is \(p=0.05\), a small sample often contains zero or only a few successes. The sample proportion \(\hat{p}\) therefore piles up near zero and can have a long tail toward larger values: it is right-skewed.

Let \(X\) be the number of successes in a random sample of size \(n\). Under a suitable binomial model, \(X\) counts the successes and \(\hat{p}=X/n\). The possible values of \(\hat{p}\) are consequently separated by steps of \(1/n\). For example, with \(n=20\), the possible proportions are \(0,\ 0.05,\ 0.10,\) and so on. When \(n\) is small and \(p\) is low, zero and a few small counts can be common, while larger counts stretch the distribution to the right.

Definition: A sampling distribution is right-skewed when it has a longer tail toward larger values. For a small population proportion, the sampling distribution of \(\hat{p}\) can be right-skewed because many samples have few or no successes, but a smaller number have noticeably more.

The mean of the sampling distribution remains \(p\), and its standard deviation is \(\sqrt{p(1-p)/n}\) under the conditions discussed earlier. Those facts describe its center and spread, not its shape. A distribution can have the correct mean and standard deviation and still be too skewed for a Normal model to be a good approximation.

Small Sample: Many Proportions Near Zero

Consider \(p=0.05\) and \(n=20\). The expected number of successes is \(np=1\), while the expected number of failures is \(n(1-p)=19\). The Large Counts condition is not met because the expected success count is below 10. We can see what that shortage means for shape by considering the binomial count \(X\).

For a binomial random variable, the probability of exactly \(x\) successes is \(\binom{n}{x}p^x(1-p)^{n-x}\). Since \(\hat{p}=X/20\), probabilities for counts translate directly into probabilities for sample proportions.

Worked Example: A Right-Skewed Distribution for a Small Sample

A quality inspector samples 20 items from a process in which 5% of items have a particular flaw. Assume the sample is random and the binomial model is appropriate. Let \(X\) be the number of flawed items and \(\hat{p}=X/20\). Find the probabilities of zero, one, and two flawed items, then describe what they suggest about shape.

Plan. Use the binomial probability formula with \(n=20\) and \(p=0.05\). The corresponding values of \(\hat{p}\) are \(0,\ 0.05,\) and \(0.10\).

Do: zero flawed items.

$$ P(X=0)=\binom{20}{0}(0.05)^0(0.95)^{20} =(0.95)^{20}\approx0.3585 $$

Do: one flawed item.

$$ P(X=1)=\binom{20}{1}(0.05)(0.95)^{19} \approx0.3774 $$

Do: two flawed items.

$$ P(X=2)=\binom{20}{2}(0.05)^2(0.95)^{18} \approx0.1887 $$

The probability of three or more flaws is the remaining probability: \(P(X\geq3)=1-[P(X=0)+P(X=1)+P(X=2)]\approx0.0755\), rounded using the unrounded probabilities above. Thus, about 35.85% of samples have \(\hat{p}=0\), and about 37.74% have \(\hat{p}=0.05\). The distribution is concentrated at zero and small proportions, with a tail extending toward larger counts. That is a strong right-skewed shape, not a bell-shaped one.

Conclude. Even though the sampling distribution is centered at \(p=0.05\), the many low values and thinner tail to the right make a Normal model inappropriate here. This matches the failed Large Counts condition: the expected number of successes is only 1.

What Changes When the Sample Gets Larger?

Keep \(p=0.05\), but increase the sample size. A larger \(n\) makes both expected counts larger, and it gives the sample proportion more possible values because its increments are \(1/n\). The probability no successes occur also falls sharply: \(P(X=0)=(1-p)^n=0.95^n\). As samples become more likely to include a range of success counts, the sampling distribution fills out around its center instead of piling up at zero.

The next example compares \(n=20\) with \(n=400\). The point is not that the larger distribution becomes perfectly symmetric; \(\hat{p}\) is still discrete and cannot be negative. Rather, its shape is much more nearly symmetric, with a Normal model better supported by the Large Counts condition.

Worked Example: Comparing a Small and a Large Sample

A monitoring team studies a characteristic found in 5% of items. Compare the sampling distributions of \(\hat{p}\) for random samples of size 20 and 400, assuming the sampling process makes the observations independent.

State. For both sample sizes, \(p=0.05\). The question is how the shape and spread of \(\hat{p}\) change as \(n\) increases.

Plan. Compare expected successes and failures to check the Large Counts condition, then calculate the standard deviation of \(\hat{p}\) and the probability of no successes. These quantities help explain the shapes; the standard deviation alone does not determine shape.

Do: sample size 20. The expected counts are \(20(0.05)=1\) success and \(20(0.95)=19\) failures. Since the expected success count is below 10, the Large Counts condition fails. The standard deviation is:

$$ \sigma_{\hat{p}}=\sqrt{\frac{0.05(0.95)}{20}} =\sqrt{0.002375}\approx0.04873 $$

The probability of no successes is \(0.95^{20}\approx0.3585\), as calculated in the previous example. A large mass at zero contributes to strong right skew.

Do: sample size 400. The expected counts are \(400(0.05)=20\) successes and \(400(0.95)=380\) failures. Both are at least 10, so the Large Counts condition is met. The standard deviation is:

$$ \sigma_{\hat{p}}=\sqrt{\frac{0.05(0.95)}{400}} =\sqrt{0.00011875}\approx0.01090 $$

The probability of no successes is \(0.95^{400}\approx0.00000000123\), or about \(1.23\times10^{-9}\). Thus, a sample proportion of zero is extraordinarily unlikely in the larger sample. Its possible values also occur in smaller steps of \(1/400=0.0025\), rather than \(1/20=0.05\).

Conclude. For \(n=400\), the sampling distribution is centered at \(0.05\), has less spread, and is much more nearly symmetric than for \(n=20\). A Normal model is supported by the Large Counts condition, assuming the sampling conditions are appropriate. The condition supports the approximation; it does not make the actual distribution exactly Normal.

The Change in Shape Is Gradual, Not a Switch

The Large Counts threshold is a practical rule, not a boundary where skewness suddenly disappears. For \(p=0.05\), the success-count requirement \(np\geq10\) requires \(n\geq200\). The failure-count requirement is much less demanding because \(1-p=0.95\). At \(n=200\), the expected counts are exactly 10 successes and 190 failures, so the Large Counts condition is met. But passing a rule of thumb does not claim perfect symmetry, nor does it guarantee that every Normal approximation will be equally accurate.

Compare \(n=199\) and \(n=200\). At \(n=199\), the expected success count is \(199(0.05)=9.95\), so the condition is not met. At \(n=200\), it is \(200(0.05)=10\), and the failure count is \(200(0.95)=190\), so both parts pass. The change in the actual distribution from 199 to 200 observations is small; only the rule’s pass-or-fail status changes at that point. Larger samples beyond this threshold generally make the approximation more convincing.

Worked Example: Interpreting the Threshold for \(p=0.05\)

A researcher wants to use a Normal model for the sample proportion of a characteristic that occurs in 5% of a population. Find the smallest sample size that meets the Large Counts condition, and explain what that result says about shape.

State and plan. Here \(p=0.05\) and \(1-p=0.95\). Both expected counts must be at least 10, so solve the two inequalities and use the larger required sample size.

Do. The success requirement is \(0.05n\geq10\), which gives \(n\geq200\). The failure requirement is \(0.95n\geq10\), which gives \(n\geq10/0.95\approx10.53\), or at least 11 as a whole-number sample size. The success requirement is more demanding, so the smallest candidate is \(n=200\).

Verify at 200:

$$ np=200(0.05)=10 \qquad\text{and}\qquad n(1-p)=200(0.95)=190 $$

Both expected counts are at least 10. At one less, \(n=199\), the expected success count is \(199(0.05)=9.95\), which is below 10. Therefore, 200 is the smallest whole-number sample size meeting the condition.

Conclude. At \(n=200\), the Large Counts condition supports a Normal model for \(\hat{p}\), provided the sampling process is appropriate. This is evidence that the sampling distribution is sufficiently close to Normal for the usual approximation; it is not proof that the distribution has become exactly symmetric at precisely \(n=200\).

Near Zero and Near One: Opposite Directions of Skew

The same reasoning applies when \(p\) is near 1. In that case, successes are common and failures are rare. The sample proportion of successes may pile up near 1, with a thinner tail extending toward smaller proportions; this is left skew. For example, if \(p=0.95\) and \(n=20\), the expected number of failures is \(20(0.05)=1\). Looking at the failure count makes the shape easier to understand: most samples have few failures, while some have more.

When \(p=0.05\), it is the successes that are rare, producing a tail toward larger sample proportions. When \(p=0.95\), it is the failures that are rare, producing a tail toward smaller sample proportions. In either case, increasing \(n\) increases the expected count of the rarer outcome and supports a more nearly symmetric sampling distribution.

Key idea: When \(p\) is near 0, the sampling distribution of \(\hat{p}\) can be right-skewed; when \(p\) is near 1, it can be left-skewed. Larger samples increase the expected count of the rarer outcome and generally make the distribution more nearly symmetric.

Common Mistakes and AP Exam Communication

  • Assuming the mean tells the whole story. The mean of \(\hat{p}\) is \(p\), but that does not imply a symmetric shape. Describe shape separately from center and spread.
  • Calling a small-\(n\), small-\(p\) distribution symmetric because it is centered at \(p\). With \(p=0.05\) and \(n=20\), the possible proportions are bounded below by zero and many samples have zero successes. Explain the resulting right tail.
  • Mixing up the sample data and the sampling distribution. The shape being discussed is the distribution of \(\hat{p}\) across repeated samples of the same size, not the distribution of individual successes and failures in one sample.
  • Claiming a larger sample makes the distribution exactly Normal. Say that it becomes more nearly symmetric and that the Large Counts condition supports a Normal approximation. The sample proportion remains discrete and bounded.
  • Treating the threshold as an abrupt shape change. The Large Counts condition is a decision rule. The distribution changes gradually as \(n\) grows, even though the rule changes from failing to passing at a particular sample size.
  • Checking only expected successes. State both \(np\) and \(n(1-p)\). The rarer outcome usually controls how large \(n\) must be.
AP Exam Tip: A strong response links the condition to the shape: “For \(p=0.05\) and \(n=20\), \(np=1\), so the expected number of successes is small. Many sample proportions will be near zero, with a tail toward larger values; the sampling distribution is right-skewed, and a Normal model is not supported by the Large Counts condition.” For a larger sample, report both expected counts and say the condition supports a more nearly symmetric Normal approximation.

Key Takeaway

When the population proportion is near zero and the sample is small, the sampling distribution of \(\hat{p}\) can cluster near zero and stretch to the right. Increasing the sample size raises the expected number of rare outcomes and generally makes the distribution more nearly symmetric. The Large Counts condition helps assess whether a Normal model is reasonable, but it does not guarantee exact symmetry.

Key takeaway: For \(p=0.05\), small \(n\) can produce a strongly right-skewed sampling distribution because few successes are expected. Larger \(n\) makes the distribution more nearly symmetric as both expected counts grow.

Check Your Understanding

Use the relationship between rare outcomes, expected counts, and shape to answer each question.

  1. For \(p=0.05\) and \(n=40\), calculate both expected counts. Is a Normal model supported by the Large Counts condition? Describe the likely shape.
  2. For \(p=0.05\) and \(n=300\), calculate both expected counts. How does the shape compare with the sampling distribution for \(n=20\)?
  3. Explain why a sample proportion of zero is much more plausible for \(p=0.05,\ n=20\) than for \(p=0.05,\ n=400\).
  4. If \(p=0.96\) and \(n\) is small, which outcome is rare, and in which direction is the sampling distribution of \(\hat{p}\) likely to have a tail?
  5. In one or two sentences, distinguish “the Large Counts condition is met” from “the sampling distribution is exactly Normal.”