From the Center to the Spread
In Mean of the Sampling Distribution of p-hat, you learned that the sampling distribution of \(\hat{p}\) is centered at the population proportion \(p\). That tells us where sample proportions tend to cluster over repeated random samples, but not how much they vary. This tutorial focuses on that spread.
The standard deviation of the sampling distribution of \(\hat{p}\) describes the typical distance between a sample proportion and its mean \(p\), across repeated samples of the same size. It is also called the standard error of \(\hat{p}\) when emphasizing that it measures the variability of a statistic from sample to sample. The formula uses the population proportion and the sample size.
The Formula and What Its Inputs Mean
For a random sample of size \(n\) from a population with proportion \(p\), the standard deviation of the sampling distribution of \(\hat{p}\) is found by multiplying \(p\) by the proportion without the characteristic, \(1-p\), dividing by \(n\), and taking the square root.
The formula can be understood by connecting \(\hat{p}\) to the number of successes in a sample. If \(X\) is the number of sampled individuals with the characteristic, then \(\hat{p}=X/n\). Under an independent-trials model, the variance of \(X\) is \(np(1-p)\). Dividing a count by \(n\) converts it to a proportion; the resulting standard deviation is \(\sqrt{np(1-p)}/n\), which simplifies to the formula above.
The formula’s answer is a proportion, such as 0.0455. To express the same spread in percentage points, multiply by 100: 0.0455 corresponds to about 4.55 percentage points. This does not mean that every sample proportion will be within 4.55 percentage points of \(p\); it describes the standard deviation of the sampling distribution.
Notice that \(p\) is the population proportion, not the sample proportion from one particular sample. Once \(p\) and \(n\) are given, the formula gives the model-based standard deviation of \(\hat{p}\). If the population proportion is not known, this exact value cannot be calculated from the formula without an assumed or specified value for \(p\).
Conditions for Using the Formula
The standard deviation formula relies on a random process that gives observations the stated success probability \(p\), and on a fixed sample size. If observations come from independent trials, the formula applies directly. If a simple random sample is drawn without replacement from a finite population, observations are not exactly independent. The 10% condition lets us treat them as approximately independent for this calculation: the sample size should be no more than 10% of the population.
The Large Counts condition, \(np\geq10\) and \(n(1-p)\geq10\), is used when deciding whether a normal model for the sampling distribution is appropriate. It is not required just to calculate the standard deviation with this formula. As in Normal Models for Sample Means and Proportions, conditions for a normal approximation address the shape of a sampling distribution, not the arithmetic of its standard deviation.
Worked Example: Calculate the Standard Deviation When \(p=0.40\) and \(n=116\)
Worked Example: Calculate the Standard Deviation When \(p=0.40\) and \(n=116\)
A community has 1,800 registered residents, and 40% of them use a local public library at least once a month. A researcher takes a simple random sample of 116 residents. Find and interpret the standard deviation of the sampling distribution of the sample proportion who use the library monthly.
Let \(\hat{p}\) be the proportion of sampled residents who use the library monthly. The population proportion is \(p=0.40\), and \(n=116\). We want \(\sigma_{\hat{p}}\).
The sample is described as a simple random sample, so the random-sampling condition is met. The 10% condition requires \(116\leq0.10(1800)=180\), which is true. The formula is appropriate for the standard deviation.
Substitute \(p=0.40\) and \(n=116\). Keep \(1-p=0.60\) in the calculation:
As a check, \(0.0455^2\approx0.002070\), which is close to \(0.24/116\approx0.002069\); the small difference comes from rounding. In percentage points, \(0.0455(100)\approx4.55\).
Across repeated simple random samples of 116 residents, the sample proportion who use the library monthly typically differs from its mean of 0.40 by about 0.0455, or 4.55 percentage points, as measured by the standard deviation.
The mean of this sampling distribution is 0.40, as established in the earlier tutorial on the mean of the sampling distribution of \(\hat{p}\). The standard deviation, about 0.0455, describes the spread around that center. It is not the standard deviation of library-use responses for individual residents.
Worked Example: A Proportion of \(0.28\)
Worked Example: A Proportion of \(0.28\)
A city’s tree inventory indicates that 28% of its street trees are young trees. A simple random sample of 80 trees is selected from a population of 1,200 street trees. Find the standard deviation of the sampling distribution of the proportion of sampled trees that are young.
Check the sampling conditions. The sample is random, and its size is fixed at 80. For the 10% condition, \(0.10(1200)=120\), and \(80\leq120\). Thus, treating the observations as approximately independent is reasonable.
Substitute and calculate. Here, \(p=0.28\), so \(1-p=0.72\):
A check gives \(0.0502^2\approx0.002520\), consistent with \(0.2016/80=0.00252\). The standard deviation is about 0.0502, or 5.02 percentage points. In repeated random samples of 80 street trees, the sample proportion that are young typically varies about 0.0502 from the population proportion 0.28, measured by the standard deviation.
Worked Example: A Proportion of \(0.72\)
Worked Example: A Proportion of \(0.72\)
A regional food bank estimates that 72% of its delivery requests arrive during a particular four-hour time window. Suppose 150 requests are randomly selected from a set of 2,500 requests, and each request is classified as either arriving during that window or not. Find the standard deviation of the sampling distribution of the proportion that arrive during the window.
Check the conditions. The requests are randomly selected and the sample size is fixed. For the 10% condition, \(0.10(2500)=250\), and \(150\leq250\). The formula is appropriate.
Calculate. Use \(p=0.72\), \(1-p=0.28\), and \(n=150\):
To check the result, \(0.0367^2\approx0.001347\), which is close to \(0.001344\); the difference reflects rounding the standard deviation to four decimal places. The standard deviation is about 0.0367, or 3.67 percentage points. It describes the spread of the sample proportions across repeated random samples of 150 requests, not the spread of individual arrival times.
Interpreting and Checking Your Result
A standard deviation must be interpreted as a measure of the sampling distribution’s spread, in the context of repeated samples. A clear interpretation names the statistic, identifies the repeated sampling process, and states the approximate distance from the mean. For example: “In repeated random samples of 116 residents, the sample proportion who use the library monthly typically differs from 0.40 by about 0.0455.”
This interpretation does not say that a randomly chosen individual resident has a 0.0455 chance of being a library user. The statistic \(\hat{p}\) describes a sample proportion; \(p\) describes the population proportion. The formula’s result concerns the variability of the statistic across samples.
The result should also be reasonable in scale. A standard deviation for a proportion is nonnegative. Its value is in proportion units and should not be mistaken for a count of people or trees. If a calculation produces a negative value, omits the square root, or uses \(p\) without its complement \(1-p\), check the setup.
Common Mistakes and AP Exam Communication
- Using \(p\) in place of \(1-p\). The formula includes both the proportion of successes and the proportion of non-successes: \(p(1-p)\). For \(p=0.40\), use \(0.40(0.60)\), not \(0.40(0.40)\).
- Forgetting the square root. The quantity \(p(1-p)/n\) is the variance of \(\hat{p}\). The standard deviation is its square root.
- Using \(\hat{p}\) instead of the stated population proportion. This formula uses the model’s \(p\). Do not substitute the proportion from one sample unless a question specifically asks for an estimate based on sample data.
- Calling \(p\) the standard deviation. The population proportion \(p\) is the mean of the sampling distribution of \(\hat{p}\); \(\sqrt{p(1-p)/n}\) is its standard deviation.
- Confusing the standard deviation of \(\hat{p}\) with variation among individuals. The formula describes how sample proportions vary over repeated samples, not how individual responses vary in the population.
- Skipping the 10% condition for a sample without replacement. State the population size and compare \(n\) with \(0.10N\). If the condition is not met, the independence approximation behind this formula needs reconsideration.
- Interpreting the result as a guaranteed distance. A standard deviation is not a maximum difference or a promise about each sample. Say “typically varies by about” or “the standard deviation is about,” not “every sample is within” that amount.
Key Takeaway
The standard deviation of the sampling distribution of \(\hat{p}\) is \(\sqrt{p(1-p)/n}\), when the stated random sampling process and independence conditions are appropriate. It describes the typical spread of sample proportions around their mean \(p\). For \(p=0.40\) and \(n=116\), the standard deviation is about 0.0455, or 4.55 percentage points.
Check Your Understanding
Use the standard deviation formula and explain what each result describes.
- A random sample of 100 people is drawn from a population with \(p=0.50\). Find \(\sigma_{\hat{p}}\).
- A random sample of 64 plants is drawn from a population with \(p=0.25\). Find the standard deviation of the sampling distribution of the proportion of plants with the characteristic.
- In Worked Example 1, why is the standard deviation interpreted as about 4.55 percentage points rather than as a standard deviation of individual residents’ responses?
- A student calculates \(p(1-p)/n\) and reports that value as the standard deviation. What step is missing?
- A simple random sample of 240 items is drawn without replacement from a population of 2,000 items. Does the 10% condition hold? Show the comparison.