Tutorials › AP Statistics › Describing a Sampling Distribution of p-hat in Context

Sampling distributions for proportions · Tutorial 418 of 1000

Describing a Sampling Distribution of p-hat in Context

Practice writing a contextual shape, center, and spread description of p-hat, with the conditions that support it.

Intermediate 9 min read

What You'll Learn

  • Describe the shape of the sampling distribution using the Large Counts condition.
  • State its center in terms of the population proportion and the context.
  • Calculate and interpret its standard deviation for a stated sample size.
  • Check the random, 10%, and Large Counts conditions explicitly.
  • Distinguish an approximately Normal shape from a skewed shape when Large Counts fails.
  • Combine shape, center, spread, conditions, and context in a free-response answer.

A Complete Description Has Three Parts

In Common Mistakes with Sampling Distributions of Proportions, you practiced keeping the population, sample, and sampling distribution distinct. Now put the distribution’s main features together. A complete description names its shape, center, and spread, and explains whether the sampling conditions support that description. The goal is not just to list numbers: connect each feature to the sample proportions that could result from repeated samples in the stated context.

Here, \(\hat{p}\) is the proportion of sampled individuals with a specified characteristic. Its sampling distribution describes the values of \(\hat{p}\) across all possible random samples of the same size from the population or process. The population proportion \(p\) and sample size \(n\) determine the center and spread; the conditions help us assess the shape.

Definition: A contextual description of the sampling distribution of \(\hat{p}\) states the shape of the distribution, its center \(p\), and its standard deviation \(\sigma_{\hat{p}}\). It also names the repeated-sample setting and checks the conditions that support the description.

Check Conditions Before Describing Shape

As covered in When the 10 Percent Condition Applies and Checking Normality of \(\hat{p}\) with \(np\) and \(n(1-p)\), conditions answer different questions. Random selection supports treating the sample as representative of the stated population or process. If sampling without replacement from a finite population of size \(N\), the 10% condition, \(n\leq0.10N\), supports treating observations as approximately independent. The Large Counts condition assesses whether a Normal model is appropriate.

Conditions: For a complete sampling-distribution description, check:
  • Random: The sample is random, or the sampling process otherwise supports treating the observations as random.
  • 10% condition: If sampling without replacement from a finite population, check \(n\leq0.10N\).
  • Large Counts condition: Check both \(np\geq10\) and \(n(1-p)\geq10\). If both hold, the sampling distribution of \(\hat{p}\) is approximately Normal in shape.

When the Large Counts condition holds, describe the shape as approximately Normal or approximately symmetric and bell-shaped. The values of \(\hat{p}\) are discrete, changing in steps of \(1/n\), so the model is an approximation rather than an assertion that the distribution is exactly continuous and Normal.

If either expected count is below 10, do not claim that the Large Counts condition supports a Normal shape. The condition’s failure does not by itself specify the exact shape. Use what the context and \(p\) indicate: when \(p\) is near 0, a small sample often produces a distribution with a cluster near 0 and a longer tail to the right, as discussed in Shape of the Sampling Distribution When \(p\) Is Near 0 or 1. Be precise about what the condition tells you, rather than labeling every distribution that fails it as having a particular shape.

State the Center and Spread in Context

For a random sample of fixed size, the mean of the sampling distribution is \(p\). This center is the long-run average of the sample proportions across repeated samples; it does not guarantee that any one sample will have \(\hat{p}=p\). The standard deviation measures the typical spread of those sample proportions around the center. Use the population proportion \(p\), not the observed \(\hat{p}\), in its formula.

$$ \mu_{\hat{p}}=p \qquad\text{and}\qquad \sigma_{\hat{p}}=\sqrt{\frac{p(1-p)}{n}} $$

Both the center and standard deviation are expressed in proportion units. For example, a center of \(0.42\) means 42%, and a standard deviation of \(0.03\) means about 3 percentage points. If you convert to percentages, convert both consistently. Avoid saying that individual people or items vary by \(\sigma_{\hat{p}}\): this describes the spread of sample proportions, not the spread of individual responses.

A strong description links every feature to its setting: “For repeated random samples of 250 households, the distribution of the sample proportion favoring the proposal is approximately Normal, centered at 0.42, with standard deviation about 0.031.” The condition checks explain why that description is appropriate.

A Reliable Free-Response Structure

Use this sequence to make your reasoning visible. It is especially useful when a question asks you to describe a sampling distribution rather than calculate a probability.

1
Name the statistic and setting.
Say what \(\hat{p}\) measures and identify the sample size and population or process.
2
Check conditions.
State the evidence that the sample is random. For sampling without replacement, compare \(n\) with \(0.10N\). Calculate \(np\) and \(n(1-p)\) for Large Counts.
3
Give center and spread.
Use \(\mu_{\hat{p}}=p\) and calculate \(\sigma_{\hat{p}}=\sqrt{p(1-p)/n}\), showing the substitution.
4
Describe shape and conclude in context.
Use the condition results to describe an approximately Normal shape or explain why that shape is not supported. Put the features together in a sentence about repeated samples.

Worked Example: A Survey About Home Composting

Worked Example: Composting in a Community

Suppose 42% of households in a town compost food scraps. A random sample of 250 households is selected without replacement from the town’s 6,000 households. Describe the shape, center, and spread of the sampling distribution of \(\hat{p}\), where \(\hat{p}\) is the proportion of sampled households that compost.

State. The population proportion is \(p=0.42\), the sample size is \(n=250\), and the population size is \(N=6{,}000\). The statistic \(\hat{p}\) is the proportion of sampled households that compost. We are describing its values across repeated random samples of 250 households.

Plan and check conditions. The sample is stated to be random. Since households are sampled without replacement, check the 10% condition: \(0.10N=0.10(6{,}000)=600\), and \(250\leq600\), so the condition is met. For Large Counts, the expected numbers of households that compost and do not compost are:

$$ np=250(0.42)=105 \qquad\text{and}\qquad n(1-p)=250(0.58)=145 $$

Both expected counts are at least 10, so the Large Counts condition is met and an approximately Normal shape is supported.

Do. The center is \(p=0.42\). Calculate the standard deviation using the stated population proportion and sample size:

$$ \sigma_{\hat{p}} =\sqrt{\frac{p(1-p)}{n}} =\sqrt{\frac{0.42(0.58)}{250}} =\sqrt{0.0009744} \approx0.0312 $$

Conclude. Across repeated random samples of 250 households from this town, the distribution of the proportion that compost is approximately Normal, centered at 0.42, with standard deviation about 0.0312, or about 3.12 percentage points. The random, 10%, and Large Counts conditions are all met.

Worked Example: When Large Counts Fails

Worked Example: A Rare Plant Disease

A nursery has 5,000 plants, and 6% are affected by a particular disease. A random sample of 100 plants is selected without replacement. Describe the shape, center, and spread of the sampling distribution of \(\hat{p}\), the proportion of sampled plants affected.

State. Here \(p=0.06\), \(n=100\), and \(N=5{,}000\). The random variable \(\hat{p}\) is the proportion of affected plants in a sample of 100. We describe its distribution across repeated samples from this nursery.

Plan and check conditions. The sample is random. The 10% condition is met because \(0.10N=0.10(5{,}000)=500\) and \(100\leq500\). The expected number of affected plants is \(np=100(0.06)=6\), which is less than 10. The expected number not affected is \(n(1-p)=100(0.94)=94\), which is at least 10. Because both counts are required, the Large Counts condition is not met; an approximately Normal shape is not supported.

Do. The center and standard deviation can still be calculated:

$$ \mu_{\hat{p}}=p=0.06 $$
$$ \sigma_{\hat{p}} =\sqrt{\frac{0.06(0.94)}{100}} =\sqrt{0.000564} \approx0.0237 $$

Because \(p\) is near 0 and the expected number of affected plants is small, the distribution is likely to have many sample proportions near 0 and a longer tail toward larger proportions. This describes a likely right-skewed pattern; it does not make the failed Large Counts condition pass.

Conclude. Across repeated random samples of 100 plants, the sampling distribution of the proportion affected is centered at 0.06, with standard deviation about 0.0237. Since the expected number of affected plants is only 6, the Large Counts condition fails and a Normal shape is not justified; the distribution is likely right-skewed because the disease is uncommon.

Worked Example: Put the Description Together

Worked Example: Support for a Library Proposal

In a city of 5,000 households, 72% support a proposed library renovation. A random sample of 160 households is selected without replacement. Write a complete description of the sampling distribution of \(\hat{p}\), the proportion in a sample who support the proposal.

State. The population proportion is \(p=0.72\), the sample size is \(n=160\), and \(N=5{,}000\). The statistic \(\hat{p}\) represents the proportion of sampled households that support the renovation.

Plan and check conditions. The sample is random. For sampling without replacement, \(0.10N=0.10(5{,}000)=500\), and \(160\leq500\), so the 10% condition is met. The expected number supporting the proposal is \(np=160(0.72)=115.2\), and the expected number not supporting it is \(n(1-p)=160(0.28)=44.8\). Both are at least 10, so the Large Counts condition is met and the distribution is approximately Normal in shape.

Do. The center is \(p=0.72\). The standard deviation is:

$$ \sigma_{\hat{p}} =\sqrt{\frac{0.72(0.28)}{160}} =\sqrt{0.00126} \approx0.0355 $$

Conclude. Across repeated random samples of 160 households from this city, the sample proportion supporting the renovation has an approximately Normal sampling distribution, centered at 0.72, with standard deviation about 0.0355, or 3.55 percentage points. The random, 10%, and Large Counts conditions support this description.

Common Mistakes and AP Exam Tips

  • Calling the shape “Normal” without checking conditions. A full-credit response calculates both \(np\) and \(n(1-p)\), and says that the distribution is approximately Normal only when both are at least 10.
  • Checking only one expected count. A large number of expected successes does not make up for too few expected failures, or vice versa. State both values and compare each with 10.
  • Reporting the observed sample proportion as the center. The center of the sampling distribution is \(p\), not the \(\hat{p}\) from one sample. Write “centered at \(p=\ldots\)” and identify the observed \(\hat{p}\) separately if the question gives one.
  • Giving a number without identifying what it measures. Say that \(\sigma_{\hat{p}}\) is the standard deviation of sample proportions across repeated samples of the stated size. It is not the standard deviation of individual responses.
  • Giving a condition without its evidence. “The 10% condition holds” is less convincing than showing \(n\leq0.10N\) with the numbers substituted. Similarly, show both expected counts for Large Counts.
  • Using “skewed” as an automatic consequence of failed Large Counts. Failure means the Normal approximation is not supported by that condition. Describe a likely direction of skew only when the context and value of \(p\) support it.
  • Leaving out the repeated-sample context. A full response says what the statistic measures, how large each sample is, and what population or process the samples come from.
AP Exam Tip: A compact, complete response includes the condition evidence and the interpretation: “The sample is random; \(n\leq0.10N\); and \(np=\ldots\) and \(n(1-p)=\ldots\), both at least 10. Thus, the sampling distribution of the proportion of [individuals] with [characteristic] in repeated samples of size \(n\) is approximately Normal, centered at \(p=\ldots\), with standard deviation \(\ldots\).”

Key Takeaway

A complete description makes the model’s shape, center, and spread meaningful in context. The center is the population proportion \(p\); the spread is \(\sqrt{p(1-p)/n}\); and the random, 10%, and Large Counts checks show what supports the description. When Large Counts fails, do not claim that the condition supports a Normal shape.

Key takeaway: Name what \(\hat{p}\) measures and the repeated-sample setting. Check the conditions with evidence, state the center and standard deviation, and describe the shape only as strongly as those checks justify.

Check Your Understanding

For each situation, plan a contextual shape, center, and spread description, and identify the relevant condition evidence.

  1. A random sample without replacement of 200 households is selected from a town of 4,000. If \(p=0.35\), check the 10% condition and both Large Counts values.
  2. For the situation in Question 1, calculate \(\sigma_{\hat{p}}\) and write a sentence interpreting it in context.
  3. A random sample of 50 items is taken from a large production process with \(p=0.08\). Calculate \(np\) and \(n(1-p)\). Is an approximately Normal shape supported by Large Counts?
  4. In Question 3, what are the center and standard deviation of the sampling distribution? Explain why the center is not determined by a particular sample’s observed \(\hat{p}\).
  5. Write a complete one- or two-sentence description for repeated random samples of size 120 when \(p=0.60\), and the sampling is without replacement from a population of 3,000. Include the condition checks and the standard deviation.