Tutorials › AP Statistics › Checking Normality of p-hat with np and n(1-p)

Sampling distributions for proportions · Tutorial 408 of 1000

Checking Normality of p-hat with np and n(1-p)

Practice checking both expected counts to decide whether a Normal model is appropriate for the sampling distribution of a sample proportion.

Intermediate 9 min read

What You'll Learn

  • Calculate the expected numbers of successes and failures using \(np\) and \(n(1-p)\).
  • Check both parts of the Large Counts condition, including cases that meet the boundary exactly.
  • Decide whether the condition supports using a Normal model for \(\hat{p}\).
  • Explain why a large expected count on one side does not compensate for a small count on the other.
  • Find the smallest whole-number sample size that meets both checks for a given population proportion.

Two Expected Counts Determine Whether a Normal Model Is Reasonable

In When the 10 Percent Condition Applies, you learned to check whether sampling without replacement creates enough dependence to affect the usual standard deviation formula. That condition and the normality check answer different questions. Now we will focus on the Large Counts condition: whether the expected numbers of successes and failures are both large enough for a Normal model of \(\hat{p}\) to be appropriate.

Here, a success means having the categorical characteristic being studied; a failure means not having it. These labels do not imply that one outcome is desirable. If the population proportion of successes is \(p\), then the expected number of successes in a sample of size \(n\) is \(np\), and the expected number of failures is \(n(1-p)\).

Condition: For the Large Counts condition for a Normal model of \(\hat{p}\), both expected counts must be at least 10: \(np\geq10\) and \(n(1-p)\geq10\). Check both inequalities. If either one fails, the condition is not met.

When both inequalities hold, the condition supports using a Normal model for the sampling distribution of \(\hat{p}\), provided the sampling process is also appropriate. As covered in Normal Models for Sample Means and Proportions, the model is centered at \(p\), with standard deviation \(\sqrt{p(1-p)/n}\) under the relevant sampling conditions. This tutorial is about checking the shape condition, not calculating probabilities from that model.

A Reliable Way to Check Both Counts

Use the population proportion \(p\), not the sample proportion \(\hat{p}\). The check describes expected counts under the population proportion or model being considered. For example, if \(p=0.30\) and \(n=50\), the expected number of successes is \(50(0.30)=15\), while the expected number of failures is \(50(0.70)=35\). These are expected counts, not a claim that every sample will contain exactly 15 successes and 35 failures.

A quick arithmetic check is that the two expected counts must add to the sample size:

$$ np+n(1-p)=n[p+(1-p)]=n $$

If the counts do not add to \(n\), revisit the complement \(1-p\) or the multiplication. Then compare each count with 10 separately. Passing one part does not make up for failing the other: the condition requires both.

Decision rule: Calculate \(np\) and \(n(1-p)\). If both are at least 10, the Large Counts condition is met and a Normal model is supported. If either is less than 10, the condition is not met, so do not rely on the Normal model based on this check.

Worked Example: Both Expected Counts Are Large Enough

Worked Example: Both Expected Counts Are Large Enough

A wildlife team takes a random sample of 80 tagged nesting sites from a region. Suppose 35% of all sites have evidence of a particular nesting feature. Let \(\hat{p}\) be the sample proportion of sites with the feature. Check whether the Large Counts condition is met.

Identify the values. The population proportion is \(p=0.35\), the sample size is \(n=80\), and the failure proportion is \(1-p=1-0.35=0.65\).

Check expected successes.

$$ np=80(0.35)=28 $$

Check expected failures.

$$ n(1-p)=80(0.65)=52 $$

The counts add to \(28+52=80\), matching the sample size. Both 28 and 52 are at least 10, so the Large Counts condition is met. A Normal model is appropriate for the sampling distribution of the sample proportion, assuming the random sampling process and any relevant independence condition are also justified. This does not mean the sample must contain exactly 28 sites with the feature; 28 is the expected count under the stated proportion.

Worked Example: Expected Successes Fall Short

Worked Example: Expected Successes Fall Short

A community garden coordinator takes a random sample of 100 garden plots to estimate the proportion with a rain barrel. Suppose \(p=0.08\) of all plots have a rain barrel. Check whether a Normal model for \(\hat{p}\) is supported by the Large Counts condition.

The failure proportion is \(1-p=0.92\). The expected count of plots with a rain barrel is:

$$ np=100(0.08)=8 $$

The expected count without a rain barrel is:

$$ n(1-p)=100(0.92)=92 $$

The counts add to \(8+92=100\), as they should. The expected failure count is at least 10, but the expected success count is only 8. Since both parts of the condition are required, the Large Counts condition is not met. We should not use a Normal model for the sampling distribution of \(\hat{p}\) based on this check, even though the sample size is 100 and the failure count is large.

Worked Example: Expected Failures Fall Short

Worked Example: Expected Failures Fall Short

A transit survey uses a random sample of 120 riders to estimate the proportion who pay with a monthly pass. Suppose \(p=0.94\) of riders use a monthly pass. Check the Large Counts condition for the sample proportion who use one.

Here, \(1-p=0.06\). The expected number of riders using a monthly pass is:

$$ np=120(0.94)=112.8 $$

The expected number not using a monthly pass is:

$$ n(1-p)=120(0.06)=7.2 $$

The expected counts add to \(112.8+7.2=120\). The expected success count is well above 10, but the expected failure count is below 10. Therefore, the Large Counts condition is not met, and a Normal model is not supported by this condition. Expected counts need not be whole numbers: they are averages under the model, so values such as 112.8 and 7.2 are valid for this check.

Worked Example: Meeting the Boundary Exactly

Worked Example: Meeting the Boundary Exactly

A school takes a random sample of 50 students to estimate the proportion who walk to school. Suppose \(p=0.20\). Check whether the Large Counts condition is met.

The expected number who walk is:

$$ np=50(0.20)=10 $$

The failure proportion is \(1-p=0.80\), so the expected number who do not walk is:

$$ n(1-p)=50(0.80)=40 $$

The counts add to \(10+40=50\). The condition uses “at least 10,” so an expected count equal to 10 passes. Both counts are at least 10; therefore, the Large Counts condition is met and a Normal model is supported, assuming the sampling conditions are also appropriate. Do not change the rule to “greater than 10.”

Finding a Sample Size That Meets the Condition

The same inequalities can help plan a sample. For a known value of \(p\), both \(np\geq10\) and \(n(1-p)\geq10\) must hold. Dividing by the positive proportions gives \(n\geq10/p\) and \(n\geq10/(1-p)\). Since the sample size must be a whole number, round up to the smallest integer that satisfies both requirements. This planning technique does not replace checking the sample’s other conditions.

Worked Example: The Minimum Sample Size for a Small Proportion

A conservation group expects that \(p=0.07\) of the wetlands in a region contain a particular plant species. What is the smallest sample size that meets the Large Counts condition?

The success-count requirement is \(n(0.07)\geq10\), so \(n\geq10/0.07\approx142.857\). The failure-count requirement is \(n(0.93)\geq10\), so \(n\geq10/0.93\approx10.753\). The first requirement is more demanding. Rounding its lower bound up gives a candidate minimum of \(n=143\).

Verify both conditions at \(n=143\):

$$ np=143(0.07)=10.01 \qquad\text{and}\qquad n(1-p)=143(0.93)=132.99 $$

Both counts are at least 10, and \(10.01+132.99=143\). To verify that 143 is the smallest possible whole-number sample size, check the preceding integer: \(142(0.07)=9.94\), which is below 10. Thus \(n=142\) fails, while \(n=143\) passes. Under this assumed proportion, 143 is the minimum sample size for the Large Counts condition.

What Passing the Check Does—and Does Not—Mean

The Large Counts condition is a practical check for whether the sampling distribution of \(\hat{p}\) is reasonably modeled by a Normal distribution. It is not a claim that the original categorical observations themselves have a Normal distribution. Each observation is still a success or a failure; the Normal model concerns the distribution of sample proportions across repeated samples.

Meeting this condition does not establish that the sample is random, that the population proportion is known without uncertainty, or that the sampling method avoids bias. If sampling is without replacement from a finite population, use the 10% condition from the earlier tutorial to assess whether the usual independence approximation is reasonable. A suitable normality check cannot fix a biased or nonrandom sample.

If either expected count is below 10, the standard AP decision is that the Large Counts condition is not met. That does not prove that every possible Normal approximation would be inaccurate by the same amount; it means the usual condition does not support using that model here. The next tutorial examines how the shape of the sampling distribution changes when \(p\) is near 0 or 1.

Common Mistakes and AP Exam Communication

  • Checking only one count. A large \(np\) does not compensate for a small \(n(1-p)\), or vice versa. A full check reports both values and compares each with 10.
  • Using \(\hat{p}\) instead of \(p\). For this check, use the population proportion or the proportion specified by the model. Do not substitute an observed sample proportion without a reason supported by the problem.
  • Forgetting the complement. The failure proportion is \(1-p\), not \(p\). Check that the two expected counts add to \(n\) as a simple arithmetic safeguard.
  • Rejecting a count equal to 10. The requirement is at least 10, so equality passes. A count of 9.9 does not pass.
  • Rounding a count before deciding. Keep the unrounded product for the comparison. For example, \(9.94\) is below 10 even though it might be casually rounded to 10.
  • Calling a Normal model guaranteed. Say that the condition is met and supports using a Normal model, provided the other sampling conditions are appropriate. The check does not remove the need to justify the sampling process.
AP Exam Tip: Show both substitutions and state the decision explicitly: “\(np=\ldots\geq10\), and \(n(1-p)=\ldots\geq10\). Both expected counts are at least 10, so the Large Counts condition is met and a Normal model for \(\hat{p}\) is appropriate, assuming the sampling conditions are satisfied.” If either value is below 10, identify which one and conclude that the condition is not met.

Key Takeaway

The Normal model for \(\hat{p}\) is supported by the Large Counts condition only when the expected numbers of successes and failures both reach the threshold. Check the two sides separately, use \(p\) and \(1-p\), and keep this shape check distinct from the 10% condition and from the question of whether the sample was selected appropriately.

Key takeaway: Calculate \(np\) and \(n(1-p)\). A Normal model is supported by the Large Counts condition when both are at least 10; if either is below 10, the condition is not met.

Check Your Understanding

For each situation, calculate both expected counts and decide whether the Large Counts condition supports a Normal model for \(\hat{p}\).

  1. A random sample of 60 devices is tested. Suppose \(p=0.25\) of all devices meet a specified battery-life standard. Check both expected counts and explain your decision.
  2. A sample of 75 residents is used to estimate the proportion who compost food waste. Suppose \(p=0.12\). Which expected count determines whether the condition passes?
  3. A sample has size \(n=40\), and the assumed population proportion is \(p=0.75\). Is the Large Counts condition met? Show both calculations, including the complement.
  4. For \(p=0.15\), find the smallest whole-number sample size that meets both parts of the Large Counts condition. Verify your answer at that sample size and at one less.
  5. In your own words, explain why a large expected number of successes cannot compensate for fewer than 10 expected failures.