Tutorials › AP Statistics › Common Mistakes with Sampling Distributions of Proportions

Sampling distributions for proportions · Tutorial 417 of 1000

Common Mistakes with Sampling Distributions of Proportions

Learn how to check the inputs and interpretation of a sampling distribution for a sample proportion, and how to correct common mistakes.

Intermediate 9 min read

What You'll Learn

  • Distinguish the population proportion p from the observed sample proportion p-hat when calculating spread.
  • Identify the roles of sample size n and population size N.
  • Explain why one sample proportion is not the sampling distribution.
  • Check the 10% condition and Large Counts condition in context.
  • Revise incorrect statements so they describe the distribution across repeated samples.

Keep the Model, the Sample, and the Population Straight

In Sampling Variability Versus Bias in Proportions, you distinguished ordinary sample-to-sample variation from bias in a sampling method. This tutorial focuses on a different source of confusion: mixing up the quantities used to describe the sampling distribution of \(\hat{p}\). Three errors appear often: using the observed \(\hat{p}\) instead of the population proportion \(p\) in the standard deviation formula, confusing sample size \(n\) with population size \(N\), and describing one sample as if it were the whole sampling distribution.

These mistakes are connected because a sampling distribution is not a description of the people in one sample. It describes the values of a statistic across all possible random samples of the same size, taken from the same population or process. The notation helps keep those ideas separate: \(p\) describes the population, \(n\) counts individuals in each sample, \(N\) counts individuals in the population, and \(\hat{p}\) is the statistic calculated from one particular sample.

Formula: For a fixed-size random sample with a population proportion \(p\), when observations can be treated as independent, the standard deviation of the sampling distribution of \(\hat{p}\) is:
$$ \sigma_{\hat{p}}=\sqrt{\frac{p(1-p)}{n}} $$
Use \(p\), not the observed \(\hat{p}\), in this formula when describing variation under a specified population model. The denominator is the sample size \(n\). For sampling without replacement, \(N\) is used to check the 10% condition, \(n\leq0.10N\).

As covered in Standard Deviation of \(\hat{p}\) Formula, this standard deviation describes the spread of sample proportions around the sampling distribution’s center. It does not describe the spread of individual responses, and it is not the difference between one sample’s \(\hat{p}\) and \(p\). The Large Counts condition, \(np\geq10\) and \(n(1-p)\geq10\), is used to assess whether a Normal model for the sampling distribution is appropriate; it is not a replacement for the standard deviation formula.

A Three-Question Check

Before calculating or interpreting a sampling distribution, ask three questions. First, what is the population proportion \(p\) specified by the situation or model? Second, how many individuals are in each sample, \(n\), and, if sampling without replacement, how large is the population \(N\)? Third, is the statement about one observed sample or about the distribution of \(\hat{p}\) across repeated samples?

1
Identify the parameter and statistic.
Use \(p\) for the population proportion in the stated model and \(\hat{p}=x/n\) for the proportion in one sample.
2
Match each size to its job.
Use \(n\) in \(\sigma_{\hat{p}}=\sqrt{p(1-p)/n}\). Use \(N\) with \(n\) to check the 10% condition when sampling without replacement.
3
Describe the right distribution.
Name the statistic \(\hat{p}\), its center and spread, and the repeated-sample setting. Keep the observed value from one sample separate.

Worked Example: Do Not Substitute the Observed Proportion

Worked Example: A Sample of Gardeners

Suppose 35% of gardeners in a large county grow vegetables at home. A random sample of 400 gardeners is taken without replacement from a county population of 8,000 gardeners. In the sample, 152 grow vegetables. Find the standard deviation of the sampling distribution of \(\hat{p}\), and identify two incorrect substitutions a student might make.

State. The stated population proportion is \(p=0.35\). The sample size is \(n=400\), the population size is \(N=8{,}000\), and the observed sample proportion is \(\hat{p}=152/400=0.38\). We want the spread of sample proportions under the model with \(p=0.35\), not a measure based on this one observed result.

Plan and check conditions. The sample is random. Since it is drawn without replacement, check the 10% condition: \(400\leq0.10(8{,}000)=800\), so it is met. For a Normal model, the expected numbers of successes and failures are \(np=400(0.35)=140\) and \(n(1-p)=400(0.65)=260\), both at least 10. The standard deviation formula uses \(p\) and \(n\).

Do. Substitute the population proportion and sample size:

$$ \sigma_{\hat{p}} =\sqrt{\frac{p(1-p)}{n}} =\sqrt{\frac{0.35(0.65)}{400}} =\sqrt{0.00056875} \approx0.02385 $$

One error would be using the observed \(\hat{p}=0.38\) in place of \(p\), producing \(\sqrt{0.38(0.62)/400}\approx0.02427\). That calculation does not describe the sampling distribution under the stated population model. Another error would be using \(N=8{,}000\) as the denominator, giving \(\sqrt{0.35(0.65)/8{,}000}\approx0.00533\). That is not the formula for this sampling distribution; the sample size \(n=400\) belongs in the denominator.

Conclude. Under the stated random-sampling model, sample proportions from samples of 400 gardeners have a standard deviation of about 0.02385. The observed \(\hat{p}=0.38\) is one result; it is not an input needed to calculate this model’s standard deviation.

Worked Example: One Sample Is Not the Distribution

Worked Example: A Rare Garden Pest

Suppose 8% of plants in a nursery are affected by a pest. A random sample of 80 plants is selected without replacement from 2,000 plants, and 8 sampled plants are affected. Describe the observed sample proportion and the sampling distribution’s mean and standard deviation. Explain whether a Normal model is supported.

State. The population proportion is \(p=0.08\), and the sample size is \(n=80\). The observed sample proportion is \(\hat{p}=8/80=0.10\). These quantities answer different questions: \(\hat{p}\) describes this sample, while the sampling distribution describes possible \(\hat{p}\) values across repeated samples of 80 plants.

Plan and check conditions. The sample is random. For sampling without replacement, \(80\leq0.10(2{,}000)=200\), so the 10% condition is met. The mean of the sampling distribution is \(p=0.08\), as established in Mean of the Sampling Distribution of \(\hat{p}\). Calculate the standard deviation using \(p\) and \(n\). For a Normal model, check both expected counts.

Do. The standard deviation is:

$$ \sigma_{\hat{p}} =\sqrt{\frac{0.08(0.92)}{80}} =\sqrt{0.00092} \approx0.03033 $$

The expected number of affected plants is \(np=80(0.08)=6.4\), which is less than 10. The expected number not affected is \(n(1-p)=80(0.92)=73.6\). Because the first expected count is below 10, the Large Counts condition is not met, so a Normal model is not supported by this condition.

Conclude. The observed proportion in this sample is 0.10. For repeated random samples of 80 plants under the stated model, the sampling distribution of \(\hat{p}\) has mean 0.08 and standard deviation about 0.03033. Do not say that this one sample “has a mean of 0.08”; that describes the sampling distribution, not the observed sample. Also do not use a Normal model for probabilities here based on the Large Counts condition.

Worked Example: Describe Repeated Samples, Not Just One Result

Worked Example: Repeated Samples of Water Bottles

Suppose 60% of bottles produced at a plant pass a particular seal check. Imagine taking random samples without replacement from a production lot of 12,000 bottles. Compare the sampling distributions of \(\hat{p}\) for sample sizes 100 and 400. Include their centers, standard deviations, and whether a Normal model is supported.

State. The population proportion is \(p=0.60\), and the two sample sizes are \(n=100\) and \(n=400\). Each \(\hat{p}\) is the proportion of bottles passing the check in one sample of the stated size. We will compare the distributions across repeated samples of each size.

Plan and check conditions. The samples are random. For \(n=100\), the 10% condition is \(100\leq0.10(12{,}000)=1{,}200\). For \(n=400\), it is \(400\leq1{,}200\). Both are met. The Large Counts values are 60 successes and 40 failures for \(n=100\), and 240 successes and 160 failures for \(n=400\); all are at least 10. Thus, a Normal model is supported for both sampling distributions.

Do. Both sampling distributions have mean \(p=0.60\). Their standard deviations are:

$$ \sigma_{\hat{p},\,n=100} =\sqrt{\frac{0.60(0.40)}{100}} =\sqrt{0.0024} \approx0.04899 $$
$$ \sigma_{\hat{p},\,n=400} =\sqrt{\frac{0.60(0.40)}{400}} =\sqrt{0.0006} \approx0.02449 $$

Conclude. Across repeated random samples of 100 bottles, the sample proportions are centered at 0.60 with standard deviation about 0.04899. Across repeated samples of 400 bottles, they are also centered at 0.60, but their standard deviation is about 0.02449. The larger samples produce less variability. A statement such as “the sample proportion is 0.60” would be misleading: 0.60 is the center of each sampling distribution, not a guarantee that every sample gives exactly 0.60.

Common Mistakes and How to Correct Them

  • Putting \(\hat{p}\) into the standard deviation formula. When \(p\) is given as the population proportion or model value, use that \(p\). The observed \(\hat{p}\) is a result from one sample. Do not swap it into a formula that describes variability under the stated model.
  • Using \(N\) instead of \(n\) in the denominator. The formula uses the number sampled, \(n\), because it describes how sample proportions vary for samples of that size. The population size \(N\) helps check the 10% condition when sampling without replacement.
  • Calling one sample the sampling distribution. A sample produces one observed \(\hat{p}\). The sampling distribution consists of the \(\hat{p}\) values from all possible samples of the same size under the specified process.
  • Giving the distribution’s center as the observed value. Under the appropriate random-sampling model, the center is \(p\). A particular sample’s \(\hat{p}\) may differ from \(p\); that difference is not an error by itself.
  • Using a Normal model after checking only the standard deviation. A standard deviation can be calculated even when the Large Counts condition fails. Check \(np\geq10\) and \(n(1-p)\geq10\) separately before using a Normal approximation.
  • Leaving the repeated-sample setting unstated. Specify what \(\hat{p}\) measures and the sample size. Saying “the standard deviation is 0.03” is incomplete unless it is clear that this is the spread of sample proportions across repeated samples of a particular size.
AP Exam Tip: Make the distinction visible in your wording: “The observed sample proportion is \(\hat{p}=\ldots\). For repeated random samples of size \(n=\ldots\), the sampling distribution of \(\hat{p}\) has center \(p=\ldots\) and standard deviation \(\ldots\).” Then state whether the 10% condition and Large Counts condition apply.

Key Takeaway

The formula and interpretation stay aligned when you keep the population, sample, and distribution separate. Use \(p\) and \(n\) to calculate \(\sigma_{\hat{p}}\); use \(N\) to check the 10% condition for sampling without replacement. Treat the observed \(\hat{p}\) as one sample result, and describe the sampling distribution as the pattern of results across repeated samples of the same size.

Key takeaway: Before reporting a sampling distribution, check: Is the standard deviation based on \(p\) and \(n\)? Is \(N\) used only where it belongs? Does the interpretation describe repeated sample proportions rather than one sample?

Check Your Understanding

For each question, identify the quantities and explain what the statement should describe.

  1. A random sample of 250 residents is drawn from a population of 5,000. Which quantity is used in the denominator of \(\sigma_{\hat{p}}\), and what role does the other population-size quantity play in checking conditions?
  2. A model specifies \(p=0.42\), and one sample has \(\hat{p}=0.47\). Which proportion belongs in the standard deviation formula when describing the sampling distribution under that model? Explain.
  3. A student says, “The sampling distribution has \(\hat{p}=0.47\) because that was the result in my sample.” Explain the difference between the observed sample proportion and the sampling distribution.
  4. For \(p=0.20\) and \(n=40\), calculate \(np\) and \(n(1-p)\). Is the Large Counts condition met?
  5. Write one sentence describing the center and spread of the sampling distribution of \(\hat{p}\) for repeated random samples of size 250 from a population with proportion \(p=0.42\).