From a Known Population Percentage to a Poll Prediction
A poll reports a sample percentage, but that result can vary from one random sample to another. If the population proportion is known, we can use the sampling distribution of \(\hat{p}\) to estimate how likely a particular poll result would be. This tutorial puts together the ideas from Normal Model for Sample Proportion Probabilities, Using normalcdf for Probabilities About \(\hat{p}\), and Describing a Sampling Distribution of \(\hat{p}\) in Context.
The essential steps are to state the event, check whether a Normal approximation is appropriate, build the model using the known population proportion and sample size, and find the area in the tail that matches the event. The final probability describes what the model predicts for repeated random samples; it does not guarantee what any one poll will report.
Translate the Poll Question into an Event
First define \(\hat{p}\) in context. For example, if the poll asks whether residents support a proposal, \(\hat{p}\) might be the proportion of sampled residents who support it. Then translate the requested result into an inequality. “At least 55%” means \(\hat{p}\geq0.55\), so we need an upper-tail probability. “At most 55%” means \(\hat{p}\leq0.55\), so we need a lower-tail probability.
The word or more points to the upper tail; or fewer points to the lower tail. Writing the event before using a calculator is a useful safeguard, as emphasized in Common Errors with Normal Calculations. Do not choose a tail just by looking at whether the cutoff itself is a high or low number: compare it with the population proportion and follow the event wording.
Check Conditions, Then Choose the Normal Area
As in earlier tutorials on sampling distributions of proportions, check the random sampling process, the 10% condition when sampling without replacement from a finite population, and the Large Counts condition. Each check addresses a different part of whether the model fits the situation.
- Random: The poll uses a random sample, or the sampling process otherwise supports treating observations as random.
- 10% condition: If sampling without replacement from a finite population of size \(N\), check \(n\leq0.10N\).
- Large Counts condition: Using the known population proportion, check both \(np\geq10\) and \(n(1-p)\geq10\).
If these conditions support the approximation, use the Normal model for \(\hat{p}\). For “at least \(c\),” calculate the area to the right of \(c\); for “at most \(c\),” calculate the area to the left. On a calculator, these are \(\operatorname{normalcdf}(c,1\text{E}99,p,\sigma_{\hat{p}})\) and \(\operatorname{normalcdf}(-1\text{E}99,c,p,\sigma_{\hat{p}})\), respectively, where \(\sigma_{\hat{p}}=\sqrt{p(1-p)/n}\).
A poll’s sample proportion changes in steps of \(1/n\), so its possible values are discrete. The Normal model is continuous and provides an approximation. In these examples, the stated cutoff corresponds to a possible sample proportion; the usual AP Statistics calculation uses the Normal area at that cutoff without a continuity correction. State that the result is approximate.
A Four-Step Plan for a Poll Probability
For a complete response, make the event and model visible, show the condition evidence, and finish with a contextual interpretation. This structure keeps a calculator result from becoming an unexplained number.
Define \(\hat{p}\), identify \(p\) and \(n\), and write the event requested by the question.
Give evidence for randomness, check \(n\leq0.10N\) if applicable, and calculate \(np\) and \(n(1-p)\).
Calculate the standard deviation, identify the matching tail, and show the Normal probability calculation.
Interpret the approximate probability as a chance for repeated samples under the model, naming the event and poll population.
Worked Example: A Poll Reports at Least 55% Support
Worked Example: Support for a Transit Plan
Suppose 50% of adults in a city support a proposed transit plan. A polling firm takes a random sample of 400 adults without replacement from the city’s 20,000 adults. What is the approximate probability that the poll finds at least 55% support?
State. Let \(\hat{p}\) be the proportion of sampled adults who support the plan. The known population proportion is \(p=0.50\), the sample size is \(n=400\), and the population size is \(N=20{,}000\). The event is \(\hat{p}\geq0.55\), so the requested area is in the upper tail.
Plan and check conditions. The adults are selected in a random sample. Because the sample is taken without replacement, check the 10% condition: \(0.10N=0.10(20{,}000)=2{,}000\), and \(400\leq2{,}000\), so the condition is met. For Large Counts:
Both expected counts are at least 10, so a Normal model for \(\hat{p}\) is appropriate.
Do. The standard deviation is:
The Normal model is approximately \(N(0.50,0.025)\). The probability matching the event is:
Conclude. If 50% of city adults support the transit plan and the polling process meets the stated conditions, there is an approximately 0.0228, or 2.28%, chance that a random sample of 400 adults will show at least 55% support. This is a model-based prediction for repeated samples, not a guarantee about a particular poll.
Worked Example: A Poll Reports at Most 55% Support
Worked Example: Interest in a Community Clinic
Suppose 60% of adults in a county would support a new community clinic. A random sample of 400 adults is selected without replacement from the county’s 18,000 adults. Estimate the probability that at most 55% of the sampled adults support the clinic.
State. Let \(\hat{p}\) be the proportion of sampled adults who support the clinic. Here \(p=0.60\), \(n=400\), and \(N=18{,}000\). “At most 55%” is the event \(\hat{p}\leq0.55\), which calls for the lower-tail area.
Plan and check conditions. The sample is random. Since sampling is without replacement, the 10% condition is met because \(0.10N=0.10(18{,}000)=1{,}800\) and \(400\leq1{,}800\). The expected counts are:
Both counts are at least 10, so the Normal approximation is supported.
Do. The standard deviation is:
Use the lower tail because the event is \(\hat{p}\leq0.55\):
Conclude. Under the model that 60% of county adults support the clinic, the probability that a random sample of 400 adults has a support proportion of at most 55% is approximately 0.0206, or 2.06%. A result this low is uncommon under the stated population proportion and sampling model.
Worked Example: Report a Tail Probability with Careful Rounding
Worked Example: Support for a School Meal Program
Suppose 60% of families in a school district support a proposed meal program. A random sample of 300 families is drawn without replacement from 12,000 families. What is the approximate probability that the poll finds at most 55% support?
State. Let \(\hat{p}\) be the proportion of sampled families who support the program. The population proportion is \(p=0.60\), \(n=300\), and \(N=12{,}000\). The event is \(\hat{p}\leq0.55\), a lower-tail event.
Plan and check conditions. The sample is random. The 10% condition holds because \(0.10N=0.10(12{,}000)=1{,}200\), and \(300\leq1{,}200\). For Large Counts:
Both counts are at least 10, so using a Normal approximation is reasonable.
Do. Calculate the standard deviation and the lower-tail probability:
Conclude. If 60% of district families support the meal program, the model estimates about a 0.0386, or 3.86%, chance that a random sample of 300 families would show at most 55% support. The probability is rounded to four decimal places; the corresponding percentage is rounded to two decimal places.
Common Mistakes and AP Exam Tips
- Using the wrong tail. A result “at least” a cutoff is an upper-tail event, while a result “at most” a cutoff is a lower-tail event. Write the inequality for \(\hat{p}\) before entering the calculator bounds.
- Using \(\hat{p}\) instead of the known \(p\) in the model. The question supplies the population proportion for the prediction. Use it as the mean and in \(\sqrt{p(1-p)/n}\); the cutoff is not a substitute for \(p\).
- Skipping a condition check. A full-credit answer gives evidence that the sample is random, shows the 10% comparison when sampling without replacement, and reports both expected counts for Large Counts.
- Confusing the standard deviation with the standard error based on sample data. For this model-based prediction with known \(p\), calculate the sampling-distribution standard deviation using the population proportion stated in the problem.
- Reporting an unexplained calculator value. Include the event, the model inputs, and the tail bounds. A calculator output alone does not show that the correct event or distribution was used.
- Rounding too early or overstating precision. Keep extra digits in the standard deviation and calculator input, then round the probability consistently. A probability such as 0.0385 is 3.85%, not 3.86%.
- Writing as though the result is certain. A probability is a chance under the stated model. Say “the model estimates an approximately 2.28% chance,” not “the poll will be within” or “the poll must be” a particular range.
Key Takeaway
A poll prediction begins with a known population proportion and a clearly stated sample size. Check the random, 10%, and Large Counts conditions, then use the corresponding Normal model and the tail that matches the event. Interpret the resulting area as an approximate chance for repeated samples under the model.
Check Your Understanding
For each situation, identify the event, check the conditions, and describe the probability calculation that matches the question.
- A random sample of 500 voters is drawn without replacement from a city of 40,000 voters. If \(p=0.48\), check the 10% condition and both Large Counts values.
- For Question 1, define \(\hat{p}\) as the proportion who support a measure. Write the event for a poll result of at least 52% support and identify the tail.
- A random sample of 250 customers is selected from a large customer population, where \(p=0.70\). Calculate the two expected counts and decide whether Large Counts supports a Normal approximation.
- For Question 3, write the Normal model’s mean and standard deviation formula. Which value should be used as the mean: \(p\) or a poll’s cutoff?
- Suppose \(p=0.55\), \(n=400\), and a random sample is drawn without replacement from a population of 15,000. Describe the event and calculator bounds for estimating the probability of a result at most 50%, and state the condition checks you would show.