Tutorials › AP Statistics › Unusual Sample Proportions and the 2 Standard Deviation Rule

Sampling distributions for proportions · Tutorial 414 of 1000

Unusual Sample Proportions and the 2 Standard Deviation Rule

Use the two-standard-deviation range of a sample proportion’s sampling distribution to assess whether an observed result is surprising under a claimed population proportion.

Intermediate 9 min read

What You'll Learn

  • Find the approximate range containing 95% of sample proportions under a claimed proportion.
  • Check randomness, the 10% condition, and the Large Counts condition before applying the rule.
  • Decide whether an observed sample proportion falls outside the two-standard-deviation range.
  • Connect the range check to the z-score for a sample proportion.
  • Explain what an unusual result suggests—and what it does not prove—about a claim.

When Is an Observed Sample Proportion Surprising?

In Computing a z-Score for a Sample Proportion, you learned to measure how far an observed \(\hat{p}\) is from a claimed population proportion \(p\), in standard deviations of the sampling distribution. This tutorial uses that distance to make a quick judgment: is the observed sample proportion unusually far from the value predicted by the claim?

The judgment depends on a model for how \(\hat{p}\) varies across repeated random samples. When the sampling distribution is approximately Normal, the empirical rule says that about 95% of its values fall within two standard deviations of its mean. For sample proportions, the model’s mean is \(p\), and its standard deviation is \(\sigma_{\hat{p}}=\sqrt{p(1-p)/n}\). So we can identify an approximate range of ordinary results and check whether the observed \(\hat{p}\) falls outside it.

Definition: Under the two-standard-deviation rule, an observed sample proportion is considered unusual if it is more than two standard deviations from the claimed population proportion in an appropriate Normal model. Equivalently, it falls outside the approximate range \(p\pm2\sigma_{\hat{p}}\).

This rule is a screening guideline based on the empirical rule, not a formal test with a universal definition of “unusual.” For a roughly Normal model, about 95% of results fall in the middle range, leaving about 5% in the two tails combined. Because the model is symmetric, about 2.5% falls in each tail. A result outside the range is therefore relatively uncommon under the model, but it can still happen by chance.

Find the Range, Then Compare

The range is centered at the claimed proportion \(p\). Its width depends on \(\sigma_{\hat{p}}\), which in turn depends on the claim and the sample size. A larger sample generally has a smaller standard deviation of \(\hat{p}\), so its two-standard-deviation range is narrower.

Formula: The approximate two-standard-deviation range for sample proportions is $$ p\pm2\sigma_{\hat{p}} =p\pm2\sqrt{\frac{p(1-p)}{n}} $$ If the observed \(\hat{p}\) is below the lower endpoint or above the upper endpoint, it is more than two standard deviations from \(p\). If it is between the endpoints, it is not more than two standard deviations away.

The z-score from the previous tutorial offers an equivalent check. Since \(z=(\hat{p}-p)/\sigma_{\hat{p}}\), an observation is outside the range precisely when \(|z|>2\). The range method makes the cutoff visible in the original proportion scale; the z-score method expresses the distance in standard-deviation units.

Before using either approach, check that the Normal model for \(\hat{p}\) is appropriate. As in Normal Model for Sample Proportion Probabilities and Computing a z-Score for a Sample Proportion, establish that the sample is random or otherwise supports treating observations as random, check the 10% condition for sampling without replacement, and verify the Large Counts condition. The empirical rule is a statement about a Normal model; it does not make an unsuitable model appropriate.

Conditions: Establish that the sample is random, or that the process supports treating observations as random. For sampling without replacement, check the 10% condition, \(n\leq0.10N\). For a Normal model of \(\hat{p}\), check the Large Counts condition: \(np\geq10\) and \(n(1-p)\geq10\).

Worked Example: A Sample Proportion Above the Range

Worked Example: Households with a Vegetable Garden

A town’s planning team claims that 30% of households have a vegetable garden. A random sample of 100 households is selected without replacement from the town’s 2,000 households. In the sample, 42 households have a vegetable garden. Use the two-standard-deviation rule to decide whether the observed sample proportion is unusual under the claim.

State. Let \(\hat{p}\) be the proportion of sampled households with a vegetable garden. The claim is \(p=0.30\), and the observed sample proportion is \(\hat{p}=42/100=0.42\).

Plan and check conditions. The sample is random. The 10% condition is met because \(100\leq0.10(2{,}000)=200\). Under the claim, the expected number with a garden is \(np=100(0.30)=30\), and the expected number without one is \(n(1-p)=100(0.70)=70\). Both are at least 10, so the Large Counts condition supports a Normal model.

Do. Calculate the standard deviation of the sampling distribution and use it to find the two-standard-deviation range:

$$ \sigma_{\hat{p}} =\sqrt{\frac{p(1-p)}{n}} =\sqrt{\frac{0.30(0.70)}{100}} =\sqrt{0.0021} \approx0.04583 $$
$$ p\pm2\sigma_{\hat{p}} =0.30\pm2(0.0458258) =0.30\pm0.09165 \approx(0.20835,\ 0.39165) $$

The observed proportion \(0.42\) is greater than the upper endpoint, about \(0.39165\). As a second check, its z-score is \(z=(0.42-0.30)/0.0458258\approx2.619\), which is greater than 2. The two checks agree.

Conclude. The sample proportion of households with a vegetable garden is more than two standard deviations above the claimed proportion of 0.30. It is unusual under the claim and the sampling model, though this result alone does not prove that the claim is false.

What “Unusual” Does and Does Not Mean

A result outside the range is evidence that the observed sample proportion does not fit comfortably with what the model typically produces. The conclusion is conditional: if the claim and the sampling model are reasonable, a result this far from \(p\) would be relatively uncommon. This is the same general reasoning used in Evaluating Claims Based on Probability: consider how surprising the observation would be if the claim were true.

Do not turn the rule into a proof. A rare result can occur, and an ordinary-looking sample result does not establish that the claim is correct. Also, an unusual result can arise because the claim is inaccurate, because the sampling process does not represent the population well, or because an assumption behind the model is not met. The calculation assesses the observation under the stated model; it cannot diagnose every possible problem with the data.

The two-standard-deviation range also treats departures in either direction as potentially unusual. If the observed \(\hat{p}\) is far below \(p\), describe it as unusually low under the claim. If it is far above \(p\), describe it as unusually high. The direction matters in the conclusion, even though the two-sided range has endpoints on both sides.

Worked Example: A Sample Proportion Inside the Range

Worked Example: Use of Refillable Bottles

A regional survey office claims that 55% of households regularly use refillable water bottles. A random sample of 200 households is selected without replacement from a population of 5,000 households. Of those sampled, 122 regularly use refillable bottles. Is the observed sample proportion unusual under the claim according to the two-standard-deviation rule?

State. Let \(\hat{p}\) be the proportion of sampled households that regularly use refillable bottles. The claimed proportion is \(p=0.55\), and the observed proportion is \(\hat{p}=122/200=0.61\).

Plan and check conditions. The sample is random. The 10% condition is met because \(200\leq0.10(5{,}000)=500\). Under the claim, the expected number of households that regularly use refillable bottles is \(np=200(0.55)=110\); the expected number that do not is \(n(1-p)=200(0.45)=90\). Both counts are at least 10, so the Large Counts condition supports a Normal model.

Do. Calculate the standard deviation and endpoints:

$$ \sigma_{\hat{p}} =\sqrt{\frac{0.55(0.45)}{200}} =\sqrt{0.0012375} \approx0.03518 $$
$$ 0.55\pm2(0.0351781) =0.55\pm0.07036 \approx(0.47964,\ 0.62036) $$

The observed proportion \(0.61\) is between \(0.47964\) and \(0.62036\). Its z-score is \(z=(0.61-0.55)/0.0351781\approx1.706\), whose absolute value is less than 2.

Conclude. The observed sample proportion of households that regularly use refillable bottles is within two standard deviations of the claimed proportion of 0.55. It is not unusually far from the claim according to this rule. That does not prove the claim is correct; it only means this sample result is not especially surprising by the two-standard-deviation guideline.

Sample Size and the Width of the Range

For a fixed claimed proportion, increasing \(n\) decreases \(\sigma_{\hat{p}}\). As a result, the two-standard-deviation range becomes narrower. This connects the unusual-result decision to How Sample Size Changes the Spread of p-hat: a difference that is ordinary for a small sample can be more than two standard deviations for a larger sample.

For example, if the claim is \(p=0.40\) and the observed proportion is \(\hat{p}=0.46\), the difference is 0.06 in either sample. With \(n=100\), the standard deviation is \(\sqrt{0.40(0.60)/100}\approx0.04899\), so the distance is \(0.06/0.04899\approx1.225\) standard deviations. With \(n=400\), the standard deviation is \(\sqrt{0.40(0.60)/400}\approx0.02449\), so the same difference is \(0.06/0.02449\approx2.450\) standard deviations. These are illustrative calculations: each actual comparison still needs its own sampling conditions checked.

The comparison shows why “a difference of six percentage points” is not enough to decide whether a sample result is surprising. The sample size affects how much \(\hat{p}\) typically varies. Always use the standard deviation for the stated \(p\) and \(n\), rather than relying on the raw difference alone.

Worked Example: An Unusually Low Sample Proportion

Worked Example: A Reported Food Allergy

A community health survey is evaluating the claim that 8% of adults in a region report a particular food allergy. A random sample of 200 adults is selected without replacement from a population of 8,000 adults. Six sampled adults report the allergy. Use the two-standard-deviation rule to assess the sample proportion.

State. Let \(\hat{p}\) be the proportion of sampled adults who report the allergy. The claim is \(p=0.08\), and the observed sample proportion is \(\hat{p}=6/200=0.03\).

Plan and check conditions. The sample is random. The 10% condition is met because \(200\leq0.10(8{,}000)=800\). Under the claim, the expected number who report the allergy is \(np=200(0.08)=16\), and the expected number who do not is \(n(1-p)=200(0.92)=184\). Both are at least 10, so the Large Counts condition supports a Normal model.

Do. Find the standard deviation and the endpoints of the range:

$$ \sigma_{\hat{p}} =\sqrt{\frac{0.08(0.92)}{200}} =\sqrt{0.000368} \approx0.01918 $$
$$ 0.08\pm2(0.0191833) =0.08\pm0.03837 \approx(0.04163,\ 0.11837) $$

The observed proportion \(0.03\) is below the lower endpoint of about \(0.04163\). As a check, \(z=(0.03-0.08)/0.0191833\approx-2.606\), so \(|z|>2\).

Conclude. The sample proportion of adults reporting the allergy is more than two standard deviations below the claimed proportion of 0.08. It is unusually low under the claim and sampling model, but the result does not by itself establish why the sample differed from the claim.

Common Mistakes and AP Exam Communication

A strong response makes the comparison visible and explains what it means in context. State the claimed \(p\), calculate the standard deviation from \(p\) and \(n\), and show the endpoints or the z-score. Then identify where the observed \(\hat{p}\) falls relative to the two-standard-deviation cutoff.

  • Using the observed proportion in the standard deviation. For this comparison, calculate \(\sigma_{\hat{p}}\) from the claimed \(p\), not from the sample’s \(\hat{p}\).
  • Using two standard deviations without checking conditions. The empirical rule applies to an appropriate Normal model. Check randomness, the 10% condition when relevant, and both expected counts in the Large Counts condition.
  • Calling every result inside the range proof that the claim is true. Being within two standard deviations means the result is not unusually far under this rule; it does not confirm the claim.
  • Calling an outside result impossible. About 5% of results fall outside the two-standard-deviation range in a Normal model. Such a result is unusual, not impossible.
  • Ignoring direction. Say whether the observed proportion is unusually high or unusually low, not just that it is “different.”
  • Overstating what the calculation establishes. A result can be unusual under the model without proving that a claim is false. Keep the conclusion conditional on the claim and sampling model.
AP Exam Tip: A full-credit conclusion identifies the observed sample proportion, says whether it is more than two standard deviations above or below the claimed \(p\), and describes that result as unusual or not unusually far under the model. Do not claim that the rule proves a population proportion.

Key Takeaway

The two-standard-deviation rule turns the spread of a sample proportion into an approximate range for ordinary results under a claim. Use it only when the sampling conditions support a Normal model, and interpret an outside result as unusual under that model—not as proof that the claim is wrong.

Key takeaway: Find \(p\pm2\sqrt{p(1-p)/n}\), then compare the observed \(\hat{p}\) with both endpoints. An observed value outside the range is more than two standard deviations from the claim and is considered unusual by the empirical-rule guideline.

Check Your Understanding

For each question, use the claimed proportion and sample size to guide the comparison, and state conclusions in context.

  1. A random sample of 150 students is drawn without replacement from a school of 2,000 students. A claim says that 40% participate in a club, and 72 sampled students participate. Check the conditions, find the two-standard-deviation range, and decide whether the result is unusual.
  2. A model claims that 25% of customers use a store’s pickup service. A random sample of 100 customers is selected from a large customer population, and 31 use the service. Find the range and explain the decision.
  3. In your own words, explain why an observed proportion inside the two-standard-deviation range does not prove that the claimed population proportion is correct.
  4. For the same claimed \(p\) and the same difference between \(\hat{p}\) and \(p\), explain why increasing the sample size can change the unusual-result decision.
  5. What does an observed proportion more than two standard deviations below a claimed \(p\) suggest, and what does it not prove?