Two Spreads, Two Roles
In Structure of a One-Proportion z-Interval, the standard error in the interval formula was calculated using the sample proportion \(\hat{p}\). This tutorial explains why. The key distinction is between the standard deviation of the sampling distribution of \(\hat{p}\), which uses the population proportion \(p\), and the estimated standard error, which uses the sample proportion \(\hat{p}\) when \(p\) is unknown.
The standard deviation describes how much sample proportions would vary from sample to sample if the population proportion were known. In many real inference problems, \(p\) is exactly what we are trying to estimate, so we cannot calculate that standard deviation directly. Instead, we use the observed sample proportion as an estimate of \(p\), producing an estimated standard error.
The formulas look almost the same. The difference is which proportion goes into the calculation:
The first formula was developed in Standard Deviation of p-hat Formula. It requires the population proportion \(p\). The second formula is used in the one-proportion \(z\)-interval because \(p\) is unknown and \(\hat{p}\) is available from the sample. This substitution is called a plug-in estimate: put an estimate of an unknown value into a formula that depends on that value.
Both quantities describe spread in sample proportions, so they are measured in proportion units. Neither is the standard deviation of individual yes-or-no responses. And the standard error is not the distance between \(\hat{p}\) and \(p\); it estimates how much \(\hat{p}\) would vary across repeated samples.
Why Substitute the Sample Proportion?
A confidence interval is built from sample data to estimate an unknown population proportion. If we used the true \(p\) to calculate the spread, we would need to know the very parameter the interval is meant to estimate. Using \(\hat{p}\) gives a spread estimate based on the observed data, allowing the interval to be calculated without knowing \(p\).
This substitution does not make the standard error identical to the true standard deviation in every sample. If \(\hat{p}\) is above \(p\), the estimated standard error might be larger; if \(\hat{p}\) is below \(p\), it might be smaller. The difference is part of estimation. Across repeated samples, \(\hat{p}\) varies, so the estimated standard error can vary too.
As in Structure of a One-Proportion z-Interval, the estimated standard error is multiplied by \(z^*\) to calculate the margin of error. Thus, it affects the interval’s width. It is an estimate of sampling variability, not a guarantee that the interval’s width exactly reflects the spread in every possible sample.
Worked Example: Comparing the Two Formulas When \(p\) Is Known
Worked Example: A Hypothetical Seed-Germination Process
Suppose a hypothetical seed-germination process has a known success proportion \(p=0.30\). A random sample of \(n=100\) seeds includes \(x=34\) that germinate. Compare the standard deviation of \(\hat{p}\) under the known process with the standard error estimated from this sample.
Find the standard deviation using \(p\). Because this example stipulates that the population proportion is known, use \(p=0.30\):
Find the standard error using \(\hat{p}\). The observed sample proportion is \(\hat{p}=34/100=0.34\), so substitute 0.34:
The estimated standard error, about 0.04737, is slightly larger than the theoretical standard deviation, about 0.04583. In this sample, \(\hat{p}=0.34\) is above \(p=0.30\), and the plug-in calculation gives a larger spread. In an actual inference problem, we ordinarily would not know \(p=0.30\), so we would report the estimated standard error, not the theoretical standard deviation.
Worked Example: The Standard Error in a Confidence Interval
Worked Example: Residents Supporting a Trail Extension
A town has 12,000 residents. A random sample of 250 residents includes 105 who support a proposed trail extension. Calculate the estimated standard error for a one-proportion \(z\)-interval. For comparison only, also calculate the theoretical standard deviation that would apply if the unknown population proportion happened to be \(p=0.40\).
State. Let \(p\) be the proportion of all town residents who support the trail extension. The sample proportion is \(\hat{p}=105/250\). The actual value of \(p\) is unknown; \(p=0.40\) is used only as a hypothetical comparison.
Plan and check conditions. The sample is random. Since it is sampled without replacement from a finite population, check the 10% condition:
The 10% condition is met. There are 105 sampled supporters and \(250-105=145\) sampled residents who do not support the extension. Both counts are at least 10, so the Large Counts condition for the interval is met.
Do. First calculate the sample proportion and estimated standard error. Use \(\hat{p}\), because the true \(p\) is not known:
For comparison, if \(p\) were 0.40, the theoretical standard deviation would be:
Conclude in context. The interval calculation uses an estimated standard error of about 0.03122, or 3.122 percentage points, for the sample proportion of residents supporting the trail extension. The comparison value of 0.03098 is the theoretical standard deviation only under the hypothetical assumption \(p=0.40\). Since the real population proportion is unknown, that theoretical value cannot be used as the actual standard deviation for this sample’s inference.
Worked Example: Estimated Standard Errors Can Differ Across Samples
Worked Example: Two Samples from a Hypothetical Service Process
Suppose a hypothetical service process has population proportion \(p=0.40\), and consider random samples of size \(n=100\). One sample has 30 successes and another has 50. Compare each sample’s estimated standard error with the same theoretical standard deviation.
Calculate the theoretical standard deviation. The process proportion and sample size are fixed:
Calculate the first estimated standard error. In the first sample, \(\hat{p}=30/100=0.30\):
Calculate the second estimated standard error. In the second sample, \(\hat{p}=50/100=0.50\):
The theoretical standard deviation is about 0.04899 for both samples because \(p\) and \(n\) are the same. The estimated standard errors differ: one is about 0.04583 and the other is 0.05000. That is expected because each estimate uses its own observed \(\hat{p}\). The sample proportion nearest 0.50 produces the larger plug-in spread here, since the product \(\hat{p}(1-\hat{p})\) is largest at 0.50. This does not mean the second sample is closer to the population proportion or that its standard error is automatically more accurate.
Common Mistakes and AP Exam Tips
- Calling both quantities the same thing. The theoretical standard deviation uses \(p\); the interval’s estimated standard error uses \(\hat{p}\). Label which one you calculate.
- Using \(p\) in a confidence interval when it is unknown. That would require information the problem does not provide. For the one-proportion \(z\)-interval, substitute \(\hat{p}\) in the standard error formula.
- Assuming the estimated standard error must be smaller. It can be larger or smaller than the theoretical standard deviation, depending on the sample’s \(\hat{p}\).
- Confusing standard error with the margin of error. The standard error estimates spread; the margin of error is \(z^*\) multiplied by the estimated standard error.
- Describing it as the error in the sample proportion. The standard error does not tell how far this particular \(\hat{p}\) is from \(p\). It estimates the sample-to-sample spread of \(\hat{p}\).
Key Takeaway
The standard deviation of the sampling distribution of \(\hat{p}\) depends on the population proportion \(p\), while the standard error in a one-proportion confidence interval estimates that spread using the observed \(\hat{p}\). They use parallel formulas but serve different roles: one describes theoretical sampling variability when \(p\) is specified; the other makes inference possible when \(p\) is unknown.
Check Your Understanding
For each question, distinguish the population proportion from the sample proportion and explain which one belongs in the calculation.
- A process has known \(p=0.25\), and sample size \(n=160\). Write the formula for the theoretical standard deviation of \(\hat{p}\). Which proportion goes into it?
- A random sample has \(x=84\) successes out of \(n=200\). Find \(\hat{p}\), then write the formula for the estimated standard error used in a one-proportion confidence interval.
- Explain why a confidence interval for an unknown \(p\) does not use \(\sqrt{p(1-p)/n}\) with the actual \(p\).
- In a hypothetical setting, the theoretical standard deviation is 0.04 and a sample’s estimated standard error is 0.043. Is that possible? Explain.
- In one sentence, distinguish what the standard error estimates from the distance between the observed \(\hat{p}\) and the unknown \(p\).