From a Claimed Percentage to a Probability
In Unusual Sample Proportions and the 2 Standard Deviation Rule, you used the spread of the sampling distribution to decide whether an observed \(\hat{p}\) was more than two standard deviations from a claimed population proportion \(p\). Now we can make a more precise comparison: use the sampling distribution to estimate the probability of getting the observed sample proportion, or one at least as extreme, if the claim is true.
This approach is useful when a manufacturer claims a particular defect rate, a school reports a participation percentage, or a community group estimates the share of residents who support a proposal. The claim supplies the model’s population proportion \(p\). The sample size \(n\) determines how much \(\hat{p}\) varies from sample to sample. Together, these values describe the sampling distribution against which an observed result can be judged.
A small tail probability means that the observed result would not often occur under the claim and model. It is evidence that the result is unusual under those assumptions, but it does not tell us the probability that the claim itself is true. Nor does it prove that the claim is false: unusual results can occur by chance, and problems with sampling or other model assumptions can also matter.
Set Up the Sampling Distribution
As in Normal Model for Sample Proportion Probabilities, when a Normal approximation is appropriate, the sampling distribution of \(\hat{p}\) has mean \(p\) and standard deviation \(\sqrt{p(1-p)/n}\). Use the claimed \(p\), not the observed \(\hat{p}\), for both quantities. The model represents the sample proportions we would see across repeated samples if the claim and sampling process were as specified.
Before calculating an area, make the direction of concern clear. If a manufacturer is concerned that its defect rate is higher than claimed, use an upper-tail probability. If a sample result could challenge a claim by being either higher or lower, use a two-tail probability. When the question specifies a direction in advance, follow it; do not choose a direction only after seeing which tail makes the observed result look more unusual.
Check the conditions before relying on the Normal approximation. The data should come from a random sample, or a process that justifies treating observations as random. If sampling without replacement from a finite population, check the 10% condition. Finally, verify both expected counts using the claimed proportion. These checks support the model; they do not guarantee that every real-world assumption is perfect.
Worked Example: A Manufacturer’s Claimed Defect Rate
Worked Example: Defects in a Production Lot
A manufacturer claims that 4% of its sealed containers have a defective lid. A quality team randomly selects 400 containers without replacement from a production lot of 10,000. It finds 25 containers with defective lids. Is this sample result unusual if the 4% claim is true?
State. Let \(\hat{p}\) be the proportion of sampled containers with defective lids. The claim is \(p=0.04\), and the observed sample proportion is \(\hat{p}=25/400=0.0625\). Since the concern is that the defect rate may be higher than claimed, we will find the upper-tail probability \(P(\hat{p}\geq0.0625)\) under the model.
Plan and check conditions. The 400 containers are randomly selected. The 10% condition is met because \(400\leq0.10(10{,}000)=1{,}000\). Under the claim, the expected number with defective lids is \(np=400(0.04)=16\), and the expected number without defective lids is \(n(1-p)=400(0.96)=384\). Both counts are at least 10, so the Large Counts condition supports a Normal model.
Do. Calculate the standard deviation using the claim, then find the area at or above the observed proportion. The standard deviation is approximately \(0.009798\). The standardized distance is shown as a check on how far the observation is from the model’s center.
Using the Normal model, \(P(\hat{p}\geq0.0625)\) is approximately \(\operatorname{normalcdf}(0.0625,1\text{E}99,0.04,0.00979796)=0.0108\), rounded to four decimal places.
Conclude. If the true defect rate is 4% and the sampling model is appropriate, a random sample of 400 containers would have a defect proportion of at least 6.25% about 1.08% of the time. This is a small upper-tail probability, so the result is unusual under the claim and model and provides evidence that the defect rate may be higher than 4%. It does not prove that the claim is false.
Choosing One Tail or Two
The tail area must match the question being asked. In the manufacturer example, the relevant concern was an increase in the defect rate. Counting low sample proportions as equally concerning would not answer that particular question. If the question had instead asked whether the defect rate differs from 4% in either direction, then both an unusually high and an unusually low result would count as evidence against the claim.
For a two-tail comparison, measure how far the observed \(\hat{p}\) is from \(p\), then include results at least that far away on both sides. In an approximately Normal sampling distribution, the two tail areas are equal when the cutoffs are equally distant from the center. This is a probability calculation about the sample statistic under a claimed model, not a calculation of the chance that the claim is true.
Worked Example: A Sample Below a Claimed Percentage
Worked Example: Support for a Community Program
A community organization claims that 60% of residents support a proposed park program. A random sample of 200 residents is drawn without replacement from 10,000 residents, and 108 support the program. Suppose departures in either direction would matter. Is the sample proportion unusual under the 60% claim?
State. Let \(\hat{p}\) be the proportion of sampled residents who support the program. The claim is \(p=0.60\), and \(\hat{p}=108/200=0.54\). The observed proportion is 0.06 below the claim, so results at or below 0.54 and at or above 0.66 are at least as far from 0.60 in either direction.
Plan and check conditions. The sample is random. The 10% condition is met because \(200\leq0.10(10{,}000)=1{,}000\). Under the claim, the expected number who support the program is \(np=200(0.60)=120\), and the expected number who do not is \(n(1-p)=200(0.40)=80\). Both counts are at least 10, so the Large Counts condition supports a Normal model.
Do. The sampling distribution has mean 0.60 and standard deviation about 0.03464. The lower and upper cutoffs are equally distant from the claimed proportion.
As a check, the observed result is about \(0.06/0.034641\approx1.732\) standard deviations below the claimed proportion. The Normal area in one tail beyond that distance is about 0.0416, so the two equal tail areas total about \(2(0.0416)=0.0832\), or approximately 0.0833 using unrounded calculator areas.
Conclude. If 60% of residents support the program and the sampling model is appropriate, results at least as far from 60% as the observed 54% would occur about 8.33% of the time. This is not an especially small probability, so the sample is not strong evidence against the claim using this two-tail comparison. It does not establish that the 60% claim is correct.
When a Result Is Not Especially Surprising
A probability is most useful when its size is connected to a clear event. For example, “0.1056” by itself says little. “If the claimed rate is correct, a sample proportion of at least 25% would occur about 10.56% of the time” identifies the model, event, and meaning of the area. This follows the contextual interpretation approach from Interpreting Normal Probabilities in Context.
A probability such as 0.1056 is not a promise that the next sample will land above the cutoff. It describes the long-run proportion of samples that would meet the event under the assumed model. Nor is it the probability of getting exactly one particular sample proportion. The Normal model is continuous, while \(\hat{p}\) can take only certain values in a sample of fixed size. We use the model to approximate the probability of being at or beyond a cutoff.
Worked Example: A Defect Result That Is Not Unusual
Worked Example: Checking a Claimed Battery Failure Rate
A company claims that 20% of its rechargeable batteries fail a basic charging check. A random sample of 100 batteries is selected from a very large production process, and 25 fail. Is this result unusual if the claim is true and the concern is a failure rate above 20%?
State. Let \(\hat{p}\) be the proportion of sampled batteries that fail the check. The claimed proportion is \(p=0.20\), and the observed proportion is \(\hat{p}=25/100=0.25\). Because the concern is a higher failure rate, the event is \(\hat{p}\geq0.25\).
Plan and check conditions. The sample is random, and the production process is very large relative to the sample, so the observations can reasonably be treated as independent; the 10% condition is satisfied for a finite population large enough to represent that process. Under the claim, the expected number of failures is \(np=100(0.20)=20\), and the expected number of batteries that pass is \(n(1-p)=100(0.80)=80\). Both are at least 10, so the Large Counts condition supports the Normal approximation.
Do. Calculate the standard deviation and the upper-tail probability:
The probability is \(\operatorname{normalcdf}(0.25,1\text{E}99,0.20,0.04)\approx0.1056\), rounded to four decimal places.
Conclude. If the failure rate is 20%, a sample of 100 batteries would have a failure proportion of at least 25% about 10.56% of the time. This result is not especially surprising under the claim and model, and it does not provide strong evidence that the failure rate is higher than 20%.
Common Mistakes and AP Exam Communication
A strong response makes the assumptions and probability event visible. Identify the claimed \(p\), calculate the standard deviation from \(p\) and \(n\), show the tail area that matches the question, and finish with a contextual interpretation. The number alone is not a complete conclusion.
- Using the observed \(\hat{p}\) in the model’s standard deviation. When judging a sample under a claim, use the claimed \(p\) in \(\sqrt{p(1-p)/n}\).
- Choosing a tail that does not match the concern. A question about an increase calls for an upper tail. Use both tails when departures in either direction are relevant.
- Forgetting model conditions. Check randomness, the 10% condition when sampling without replacement, and both expected counts under the claim before relying on a Normal approximation.
- Calling the probability the chance that the claim is true. A tail probability is calculated assuming the claim is true; it measures how often the specified sample result or a more extreme result would occur under that assumption.
- Claiming that an ordinary result proves the claim. A result that is not unusual under the model is compatible with the claim, but it does not confirm the claim as true.
- Leaving out context and direction. Say what the sample proportion measures and whether it is unusually high, unusually low, or unusually far in either direction.
Key Takeaway
A claimed population percentage and sample size determine a sampling distribution for \(\hat{p}\). Once its conditions are checked, a Normal area estimates how often the observed result, or one at least as extreme in the relevant direction, would occur if the claim were true. Interpret that probability as evidence about the sample’s compatibility with the claim, not as proof for or against it.
Check Your Understanding
For each situation, identify the appropriate direction of concern, check the conditions, and interpret any probability in context.
- A manufacturer claims that 6% of its light bulbs fail an inspection. In a random sample of 300 bulbs from a lot of 8,000, 27 fail. If the concern is a higher failure rate, find the approximate probability of a sample proportion at least this high under the claim.
- A school claims that 45% of students take part in an arts activity. A random sample of 200 students produces \(\hat{p}=0.50\). If departures in either direction matter, identify the two cutoffs for results at least as far from the claim as this sample result.
- In your own words, explain why the probability of an observed result or one more extreme is not the probability that the population claim is true.
- Why should the claimed proportion, rather than the observed sample proportion, be used to calculate the standard deviation for this comparison?
- Describe how the conclusion changes when a question about a higher rate is replaced by one where either a higher or lower rate would matter.