Tutorials › AP Statistics › Sample Size Needed for a Desired Margin of Error

One-proportion confidence intervals · Tutorial 432 of 1000

Sample Size Needed for a Desired Margin of Error

Use the confidence level, desired margin of error, and a planning estimate of the proportion to find how many observations to collect.

Intermediate 10 min read

What You'll Learn

  • Rearrange the one-proportion margin-of-error formula to solve for sample size.
  • Use a planning proportion of 0.5 when no useful prior estimate is available.
  • Use a prior estimate and its complement to calculate a less conservative sample size.
  • Round a calculated sample size up and explain why rounding down may not meet the target.
  • Check what can be assessed about the sampling conditions before data are collected.
  • Distinguish the planned margin of error from the margin calculated using the eventual sample proportion.

Plan the Sample Before Collecting Data

In How Sample Size Affects Interval Width, we saw that a larger sample generally reduces the margin of error for a one-proportion \(z\)-interval. A practical planning question reverses that relationship: if we want a margin of error no larger than a specified amount, how large a sample should we collect?

The one-proportion \(z\)-interval margin of error is \(z^*\sqrt{\hat{p}(1-\hat{p})/n}\). Before collecting a sample, however, we do not yet know \(\hat{p}\). To plan, we substitute a value \(p^*\) as an estimate of the proportion expected in the sample. This planning value might be a prior estimate, or it might be the conservative value 0.5.

Definition: A planning proportion, written \(p^*\), is a value used in advance to estimate the standard error when selecting a sample size. It is not the unknown population proportion \(p\), and it is not the sample proportion \(\hat{p}\) from data not yet collected.

Let \(m\) be the largest desired margin of error. Replacing \(\hat{p}\) in the margin-of-error formula with the planning value \(p^*\), then solving for \(n\), gives the sample-size formula below. The sample size is the quantity being planned, so \(n\) must be a whole number at least as large as the calculated result.

$$ m=z^*\sqrt{\frac{p^*(1-p^*)}{n}} \qquad\Longrightarrow\qquad n=\frac{(z^*)^2p^*(1-p^*)}{m^2} $$

After calculating \(n\), always round up to the next whole number. Rounding to the nearest integer, or down, could leave the planned margin of error slightly larger than the target. Use the critical value \(z^*\) for the desired confidence level, as in Finding Critical Values \(z^*\) for Common Confidence Levels.

Why 0.5 Is the Conservative Choice

When there is no useful prior information about the proportion, use \(p^*=0.5\). The product \(p^*(1-p^*)\) is largest at 0.5: it equals 0.25 there, and is smaller for any other proportion between 0 and 1. Since this product determines the estimated standard error in the planning formula, using 0.5 gives the largest planned margin of error for a given sample size. It is therefore called the conservative choice for sample-size planning.

$$ 0\leq p^*(1-p^*)\leq0.25 \qquad\text{with the maximum at }p^*=0.5 $$

A conservative choice does not mean that the true population proportion is assumed to be 0.5. It means that, without a better estimate, we plan for the value that requires the largest sample to reach the target margin of error. This can lead to a larger planned sample than using a credible prior estimate, but it avoids relying on an estimate that may be inaccurate.

Worked Example: Use 0.5 for a 95% Confidence Level

A town wants to estimate the proportion of its households that compost food scraps. Planners want a 95% confidence interval with a margin of error no greater than 0.04. They have no useful prior estimate. Find the minimum sample size to plan for.

1
State.
Let \(p\) be the true proportion of households in the town that compost food scraps. The target margin of error is \(m=0.04\), the confidence level is 95%, and no prior estimate is available. Use \(p^*=0.5\).
2
Plan and check what is knowable.
For a 95% confidence level, \(z^*=1.96\). The sample will be selected randomly from the town’s 30,000 households. Once the required \(n\) is calculated, check whether it is no more than 10% of the population. The eventual Large Counts condition depends on the observed number of households that compost and do not compost; before data collection, it cannot be verified from actual counts.
3
Do.
Substitute the target margin, critical value, and planning proportion: \(n=(1.96)^2(0.5)(0.5)/(0.04)^2=0.9604/0.0016=600.25\). Round up to \(n=601\) households. The 10% condition is met because \(601\leq0.10(30{,}000)=3{,}000\). Under the planning value, the expected success and failure counts are each \(601(0.5)=300.5\), comfortably above 10. If we had rounded down to 600, the conservative planned margin of error would be \(1.96\sqrt{0.25/600}\approx0.04001\), slightly greater than 0.04. At 601, it is \(1.96\sqrt{0.25/601}\approx0.03998\).
4
Conclude.
The town should plan to obtain responses from at least 601 randomly selected households. The calculation uses the conservative planning value of 0.5, a 95% confidence level, and a target margin of error of 0.04. The number is the required completed sample size, not necessarily the number of households the town must contact if some do not respond.

Using a Prior Estimate

If a credible earlier survey or other relevant information suggests a likely proportion, planners can use it as \(p^*\). Use the estimate’s complement, \(1-p^*\), as well; do not use the estimate alone in the formula. A prior estimate generally reduces the required sample size when it is away from 0.5, because \(p^*(1-p^*)\) is then less than 0.25.

This approach depends on the prior estimate being reasonable for the population and characteristic being studied. An estimate from a different population, an outdated setting, or a biased sample may not be a sound planning value. The planned sample size is not a promise that the eventual interval will have exactly the target margin: the interval calculated after data collection uses the observed \(\hat{p}\), not the planning value.

Worked Example: Use a Prior Estimate

A regional waste program is planning a random survey to estimate the proportion of households that separate food waste for collection. A recent, relevant survey suggests that about 32% do so. Find the sample size for a 95% confidence level and a target margin of error no greater than 0.05.

Let \(p\) be the true proportion of households in the region that separate food waste. Use the prior estimate \(p^*=0.32\), so \(1-p^*=0.68\). With \(z^*=1.96\) and \(m=0.05\),

$$ n=\frac{(1.96)^2(0.32)(0.68)}{(0.05)^2} =\frac{3.8416(0.2176)}{0.0025} =\frac{0.83593216}{0.0025} =334.372864 $$

Round up to \(n=335\). Rounding down to 334 would not meet the target under the planning estimate: the estimated margin of error would be \(1.96\sqrt{(0.32)(0.68)/334}\approx0.05003\). At 335, it is \(1.96\sqrt{(0.32)(0.68)/335}\approx0.04995\). Thus, plan for at least 335 completed responses, assuming the prior estimate is suitable and the sampling conditions are supported.

For comparison, using the conservative value \(p^*=0.5\) with the same confidence level and margin of error would require \(n=384.16\), rounded up to 385. The prior estimate reduces the planned size by 50 responses in this calculation. That saving comes with a trade-off: if the prior estimate poorly reflects the current population, the realized margin of error may differ from the planning target.

What the Planning Calculation Can and Cannot Promise

The formula makes the effect of the planning choices visible. A higher confidence level uses a larger \(z^*\), so it requires a larger sample for the same margin of error. A smaller target margin also increases the required sample size; because \(m\) is squared in the denominator, cutting the target margin in half requires four times the sample size when the other inputs stay fixed.

Key relationship: For a fixed planning proportion, the required sample size is proportional to \((z^*)^2\) and inversely proportional to \(m^2\). A smaller desired margin of error can require a substantially larger sample.

The calculation is for planning a one-proportion \(z\)-interval. It does not repair problems with data collection. A large sample selected by convenience, for example, may still be biased and may not represent the population of interest. As discussed in Constructing a One-Proportion \(z\)-Interval by Hand, check the conditions for the interval that will actually be calculated.

Some checks are possible in advance. You can plan to take a random sample, and after finding \(n\), compare it with the population size \(N\) to check the 10% condition when sampling without replacement. The Large Counts condition for the actual interval uses the observed success and failure counts: \(x\geq10\) and \(n-x\geq10\). Before collecting data, the actual counts are unknown. Counts calculated from \(p^*\) are planning estimates, not proof that the condition will hold in the completed sample.

Using 0.5 has a useful property for the estimated margin of error: for any observed \(\hat{p}\) between 0 and 1, \(\hat{p}(1-\hat{p})\leq0.25\). Thus, with the planned sample size, the plug-in margin of error calculated from \(\hat{p}\) will be no greater than the conservative value, provided the one-proportion interval conditions are met. A calculation based on a prior estimate does not have that same bound; an observed proportion farther from 0.5 than expected could produce a larger margin of error than planned.

Worked Example: A Tighter Margin Requires a Larger Sample

A community garden group wants to estimate the proportion of residents who would use a shared tool library. It has no prior estimate and plans a random sample from a city of 18,000 residents. How many residents should it sample for a 95% confidence level and a margin of error no greater than 0.03?

Use the conservative planning value \(p^*=0.5\), its complement \(0.5\), and \(z^*=1.96\). The sample-size calculation is

$$ n=\frac{(1.96)^2(0.5)(0.5)}{(0.03)^2} =\frac{0.9604}{0.0009} =1067.111\ldots $$

Round up to \(n=1068\). The 10% condition is met because \(1068\leq0.10(18{,}000)=1{,}800\). Under the planning value, the expected success and failure counts are each \(1068(0.5)=534\), so the planned counts are large. The actual Large Counts condition must still be checked using the number who do and do not say they would use the tool library.

The conservative planned margin of error at \(n=1068\) is \(1.96\sqrt{0.25/1068}\approx0.02999\), which is just below 0.03. The group should plan for at least 1068 completed responses. This is about four times the sample size required for a 0.06 margin at the same confidence level and planning proportion, because halving the margin requires four times as many observations.

Common Mistakes and AP Exam Tips

  • Rounding to the nearest whole number. Sample size must be rounded up. If the formula gives 334.37, a sample of 334 is not enough to guarantee the planned margin under the stated inputs.
  • Using \(p^*\) as though it were the known population proportion. It is a planning estimate. The population parameter remains unknown, and the interval after data collection is centered at the observed \(\hat{p}\).
  • Forgetting the complement. The formula uses \(p^*(1-p^*)\). For a prior estimate of 0.32, the correct product is \((0.32)(0.68)\), not \((0.32)(0.32)\).
  • Using 0.5 because the true proportion is assumed to be 0.5. Use it when there is no suitable prior estimate because it gives the largest required sample size, not because it asserts what \(p\) must be.
  • Claiming the target margin is guaranteed in every situation. The calculation plans the margin under its inputs and assumptions. A prior estimate can be inaccurate; the actual interval uses \(\hat{p}\), and the sampling conditions must be checked.
  • Checking only the sample size and ignoring the design. Random sampling and the 10% condition still matter. A larger sample cannot remove bias from a poor sampling method.
AP Exam Tip: Show the formula with \(z^*\), \(m\), and \(p^*(1-p^*)\) identified; substitute values; and round up. Then state the planned number of observations in context. Distinguish conditions that can be checked before sampling, such as the 10% condition, from the Large Counts condition, which must be checked with observed counts for the completed sample.

Key Takeaway

To plan a one-proportion confidence interval, set the desired margin of error equal to \(z^*\sqrt{p^*(1-p^*)/n}\) and solve for \(n\). Use \(p^*=0.5\) when no useful prior estimate is available; use a credible prior estimate when appropriate, together with its complement. Always round up, and remember that the actual interval and its conditions are assessed after data collection.

Key takeaway: The planned sample size is \(n=(z^*)^2p^*(1-p^*)/m^2\). A planning proportion of 0.5 is conservative, a suitable prior estimate can reduce the required sample, and the calculated sample size must be rounded up.

Check Your Understanding

Use the sample-size planning formula and the condition distinctions in this tutorial to answer each question.

  1. For a 95% confidence level, a target margin of error of 0.05, and no prior estimate, calculate the required sample size and state how it should be rounded.
  2. A relevant prior estimate is 0.20. What value should be used for its complement in the sample-size formula?
  3. Why is 0.5 called the conservative planning proportion?
  4. A calculation gives \(n=248.02\). What sample size should be planned, and why is rounding down inappropriate?
  5. Which condition for a one-proportion interval can be checked using the planned sample size and population size? Which condition requires the observed success and failure counts?