Putting the Sampling-Distribution Ideas Together
The previous tutorial, Predicting Poll Results from a Known Population Proportion, used a known population proportion to estimate the chance of a particular poll result. In an exam-style question, that probability is often just one part of a larger response. You may also need to define the sample proportion, justify a Normal model, calculate its mean and standard deviation, and explain what the probability means.
A reliable response follows the logic of the question: identify the statistic and sampling process, check conditions, describe the sampling distribution, calculate the requested probability, and conclude in context. The mean and standard deviation are not extra details to tack on at the end; they describe the model used for the probability.
A Complete Plan for a Multi-Part Question
Before calculating, translate each part of the question into a specific task. A request to “describe the sampling distribution” calls for shape, center, and spread. A request for a probability also requires identifying the event and the matching area. As in Describing a Sampling Distribution of \(\hat{p}\) in Context, name what \(\hat{p}\) measures and describe repeated samples of the stated size.
Define \(\hat{p}\) in context, identify \(p\), \(n\), and, if given, \(N\). Write the requested result as an inequality or interval.
Use details from the sampling process to address randomness. If sampling without replacement, check the 10% condition. Calculate both expected counts for the Large Counts condition.
When the conditions support the Normal approximation, give the shape, mean, and standard deviation. Show the Normal-area setup that matches the event.
State what the probability describes for repeated samples under the model. Include the population, sample size, event, and approximate probability.
The condition checks are evidence for using the model; they are not themselves probability calculations. Keep them explicit. For sampling without replacement from a finite population, compare \(n\) with \(0.10N\). For Large Counts, use the population proportion in both \(np\) and \(n(1-p)\). If all checks support the approximation, the Normal model is approximately \(N\left(p,\sqrt{p(1-p)/n}\right)\), with \(p\) as the mean and the square-root expression as the standard deviation.
Worked Example: An Upper-Tail Probability
Worked Example: Support for a Neighborhood Garden
Suppose 42% of residents in a town support creating a neighborhood garden. A community group selects a random sample of 500 residents without replacement from the town’s 12,000 residents. Describe the sampling distribution of \(\hat{p}\), the proportion in the sample who support the garden, and estimate the probability that at least 46% support it.
State. Here \(p=0.42\), \(n=500\), and \(N=12{,}000\). The requested event is \(\hat{p}\geq0.46\), an upper-tail event.
Plan and check conditions. The problem states that the group selects a random sample, supporting the random condition. Since sampling is without replacement, check the 10% condition:
The 10% condition is met. The Large Counts values are:
Both expected counts are at least 10, so the Normal approximation is supported.
Do. The sampling distribution is approximately Normal. Its mean is \(p\), and its standard deviation is:
Thus, \(\hat{p}\approx N(0.42,0.02207)\). For the upper-tail event:
Conclude. If 42% of town residents support the garden and the stated sampling model is reasonable, there is approximately a 0.0350, or 3.50%, chance that a random sample of 500 residents will have a support proportion of at least 46%. This is a model-based chance for repeated samples, not a prediction that any one sample must produce a particular result.
Worked Example: A Lower-Tail Probability
Worked Example: Use of a Password Manager
Suppose 27% of customers of a technology service use its password manager. A random sample of 400 customers is selected without replacement from 10,000 customers. Describe the sampling distribution of the sample proportion who use the password manager and estimate the probability that at most 23% of the sample use it.
State. Let \(\hat{p}\) be the proportion of sampled customers who use the password manager. The values are \(p=0.27\), \(n=400\), and \(N=10{,}000\). The event is \(\hat{p}\leq0.23\), so we need a lower-tail area.
Plan and check conditions. The sample is stated to be random. The 10% condition holds because \(0.10N=0.10(10{,}000)=1{,}000\), and \(400\leq1{,}000\). For Large Counts:
Both expected counts are at least 10. The Normal approximation is supported.
Do. The sampling distribution is approximately Normal with mean \(0.27\). Its standard deviation is:
Using the unrounded standard deviation in the calculator setup gives:
Conclude. Under the model in which 27% of these customers use the password manager, the probability that a random sample of 400 has a sample proportion of at most 23% is approximately 0.0358, or 3.58%. The result is uncommon under the stated model, but the probability does not establish what caused a particular sample result.
Worked Example: A Probability Between Two Cutoffs
Worked Example: Recycling in an Apartment Complex
Suppose 64% of households in a large apartment complex regularly recycle. A random sample of 300 households is selected without replacement from 6,000 households. Describe the sampling distribution of \(\hat{p}\), the sample proportion that regularly recycles, and estimate the probability that \(\hat{p}\) is between 0.60 and 0.68, inclusive.
State. Here \(p=0.64\), \(n=300\), and \(N=6{,}000\). The event is \(0.60\leq\hat{p}\leq0.68\), so the probability is the area between the two cutoffs.
Plan and check conditions. The sample is random. Because it is drawn without replacement, check the 10% condition: \(0.10N=0.10(6{,}000)=600\), and \(300\leq600\). For Large Counts:
Both counts are at least 10, so using a Normal approximation is reasonable.
Do. The sampling distribution is approximately Normal with mean \(p=0.64\). Its standard deviation is:
The probability between the two cutoffs is:
Conclude. If 64% of households in the complex regularly recycle and the stated sampling conditions are reasonable, the model estimates about an 0.8511, or 85.11%, chance that a random sample of 300 households will have a recycling proportion between 0.60 and 0.68, inclusive.
Common Mistakes and AP Exam Tips
- Giving a formula without its value. A description should report the mean and the calculated standard deviation, then say what each describes. For example, the mean is the center of the sample-proportion distribution; the standard deviation describes its typical spread across repeated samples.
- Using the wrong proportion in the standard deviation. In these model-based questions, use the given population proportion \(p\), not the probability cutoff or a sample proportion, in \(\sqrt{p(1-p)/n}\).
- Stopping after naming a condition. “The 10% condition holds” is not enough by itself. Show the comparison \(n\leq0.10N\). For Large Counts, report both \(np\) and \(n(1-p)\), and verify that each is at least 10.
- Using the wrong Normal area. “At least” calls for an upper tail, “at most” calls for a lower tail, and “between” calls for the area between two bounds. Write the event before entering the calculator bounds.
- Rounding the spread too early. Keep extra digits in \(\sigma_{\hat{p}}\) for the calculator input, then round the reported probability consistently. A small rounding difference in the standard deviation can change the last digit of the probability.
- Giving a conclusion without context. A full-credit conclusion identifies the sample, the result being considered, and the probability under the model. Avoid saying the probability proves that the population proportion is different.
Key Takeaway
A strong multi-part response makes the route from conditions to conclusion easy to follow. Define \(\hat{p}\), show evidence for each condition, report the mean and standard deviation with their roles, and use the Normal area that matches the event. Finish by interpreting the probability in context as a model-based chance for repeated samples.
Check Your Understanding
For each situation, identify what you would state, check, calculate, and conclude. Use a Normal approximation only when the conditions support it.
- A random sample of 250 households is selected without replacement from a town of 8,000 households, where \(p=0.36\). Check the 10% condition and both Large Counts values.
- For Question 1, define \(\hat{p}\) as the proportion of sampled households that have solar panels. State the mean of its sampling distribution and write the standard-deviation formula using the given values.
- A random sample of 300 library members is selected from a large population where \(p=0.52\). Write the event for a sample proportion of at least 0.56 and identify the Normal tail needed.
- For Question 3, calculate the mean and standard deviation of the sampling distribution and describe what each value represents in context.
- Suppose the conditions are met and the probability of a stated event is 0.0412. Write a contextual conclusion that describes this probability as a chance for repeated samples rather than a guarantee about one sample.