One Sample, Two Different Questions
In Calculating a One-Proportion z-Test by Hand, you used \(p_0\) to calculate the standard error in a test statistic. In a one-proportion confidence interval, the standard error instead uses the observed sample proportion \(\hat{p}\). These formulas differ because a test and an interval use the data to answer different questions.
A test asks whether the sample result would be unusual if the null hypothesis were true. Its model therefore assumes \(p=p_0\). A confidence interval estimates the unknown population proportion \(p\); because \(p\) is unknown, the interval uses \(\hat{p}\) to estimate the variability of sample proportions.
What Each Standard Error Represents
If the true population proportion were \(p\), the standard deviation of the sampling distribution of \(\hat{p}\) would be \(\sqrt{p(1-p)/n}\), under the conditions for the one-proportion model. In a test of \(H_0:p=p_0\), the null hypothesis supplies the value to use for \(p\). The null model’s standard deviation is therefore \(SE_0=\sqrt{p_0(1-p_0)/n}\). It describes how much sample proportions would typically vary if the null value were correct.
For an interval, there is no single null value assumed to be true. The goal is to estimate the unknown \(p\), so the interval estimates the sampling variability by substituting \(\hat{p}\) for \(p\). This is why its estimated standard error is \(\sqrt{\hat{p}(1-\hat{p})/n}\). The two calculations may be close or noticeably different, depending on how far \(\hat{p}\) is from \(p_0\).
Using \(\hat{p}\) in a test would answer a different question. It would estimate variability around the observed result rather than model the variation expected under \(H_0\). The test statistic compares the observed \(\hat{p}\) with \(p_0\), so its denominator must come from the null model being tested.
Both expressions have the same structure: a proportion multiplied by one minus that proportion, divided by the sample size, and then square-rooted. What changes is the proportion entered. If \(\hat{p}=p_0\), the two standard errors are equal. If they differ, neither standard error is automatically larger: the product \(q(1-q)\) depends on how close \(q\) is to 0.5.
Worked Examples
Worked Example: When the Two Standard Errors Match
A fictional random sample of 150 members of a community garden finds that 90 have used the shared tool shed this season. Consider \(H_0:p=0.60\) and \(H_a:p\ne0.60\), where \(p\) is the true proportion of garden members who have used the shed. Compare the standard errors used for a test and an interval, and calculate the test statistic.
State: The parameter \(p\) is the true proportion of community garden members who have used the shared tool shed this season. The null proportion is \(p_0=0.60\). From \(x=90\) successes among \(n=150\), the sample proportion is:
Plan and check conditions: The members were randomly sampled, so the Random condition is met. If the sample was drawn without replacement from at least 1,500 members, the 10% condition holds because \(150\leq0.10(1500)=150\). For the test, the expected numbers of successes and failures under \(H_0\) are \(np_0=150(0.60)=90\) and \(n(1-p_0)=150(0.40)=60\). Both are at least 10, so the test’s Large Counts condition is met. For the interval, the observed counts are 90 successes and 60 failures; both are also at least 10.
Do: Calculate each standard error using its own proportion:
Because \(\hat{p}=p_0\), the standard errors match. The test statistic is:
Conclude: The observed sample proportion is exactly the null proportion, so it is zero null-model standard errors away from \(0.60\). The equal standard errors do not mean the procedures are interchangeable: the test evaluates a null claim, while the interval estimates \(p\).
Worked Example: The Interval Standard Error Is Larger
A fictional random sample of 100 customers at a subscription service finds that 40 use paperless billing. Test \(H_0:p=0.30\) against \(H_a:p>0.30\), where \(p\) is the true proportion of customers who use paperless billing. Compare the test and interval standard errors and calculate the test statistic.
State: The parameter is the true proportion of subscription-service customers who use paperless billing. Here, \(x=40\), \(n=100\), and \(p_0=0.30\), so:
Plan and check conditions: The customers were randomly sampled, meeting the Random condition. If sampling without replacement from at least 1,000 customers, the 10% condition holds because \(100\leq0.10(1000)=100\). Under the null, the expected counts are \(100(0.30)=30\) successes and \(100(0.70)=70\) failures, both at least 10. For the interval, the observed counts are 40 successes and 60 failures, also both at least 10.
Do: Under the null, \(p_0(1-p_0)=0.30(0.70)=0.21\). In the interval calculation, \(\hat{p}(1-\hat{p})=0.40(0.60)=0.24\). Thus:
The interval’s estimated standard error is larger because \(0.40(0.60)\) is larger than \(0.30(0.70)\). The test statistic, however, uses \(SE_0\):
Conclude: The observed sample proportion is about 2.182 null-model standard errors above 0.30. That comparison uses the null-based standard error because the test asks how far the result is from what the null model predicts. The interval standard error answers a different question by estimating variability using the observed proportion.
Worked Example: A Test Can Pass Its Count Check When an Interval Does Not
A fictional random sample of 100 households finds that 4 compost food scraps. A community planner wants to test \(H_0:p=0.20\) against \(H_a:p<0.20\), where \(p\) is the true proportion of households in the service area that compost food scraps. Compare the two standard errors and check the Large Counts condition for each procedure.
State: Here, \(x=4\), \(n=100\), and \(p_0=0.20\). The observed sample proportion is:
Plan and check conditions: The households were randomly sampled, meeting the Random condition. If they were sampled without replacement from at least 1,000 households in the service area, the 10% condition holds because \(100\leq0.10(1000)=100\). For the test, the null expected counts are \(100(0.20)=20\) successes and \(100(0.80)=80\) failures. Both meet the test’s Large Counts condition. For an interval, the observed counts are 4 successes and 96 failures. Since 4 is less than 10, the interval’s Large Counts condition is not met.
Do: The test calculation uses \(p_0=0.20\), while the interval standard-error calculation uses \(\hat{p}=0.04\):
The test statistic is:
Conclude: The sample proportion is 4.000 null-model standard errors below 0.20. The test’s null expected counts meet its Large Counts condition, but the observed success count is too small for the usual one-proportion \(z\)-interval condition. Although the interval standard-error expression can be evaluated, that arithmetic does not make a \(z\)-interval appropriate here. The two procedures check different counts because one uses the null model and the other uses the observed sample.
Common Mistakes and What Full Credit Says
The important step is not memorizing two formulas in isolation. It is identifying which model and which question the calculation belongs to. Avoid these errors:
- Using \(\hat{p}\) in a test’s standard error. This makes the denominator describe estimated variation around the sample result, not variation under \(H_0\). A full-credit test calculation shows \(SE_0=\sqrt{p_0(1-p_0)/n}\).
- Using \(p_0\) in an interval’s standard error. A confidence interval estimates the unknown population proportion and uses \(\hat{p}\) in its standard-error calculation. Do not substitute the benchmark from a separate claim.
- Assuming the two standard errors must be equal. They match when \(\hat{p}=p_0\), but can differ otherwise. Compare the actual products \(p_0(1-p_0)\) and \(\hat{p}(1-\hat{p})\) rather than guessing which is larger.
- Using the wrong counts for a Large Counts check. For a test, check \(np_0\) and \(n(1-p_0)\). For an interval, check \(n\hat{p}=x\) and \(n(1-\hat{p})=n-x\). As emphasized in the earlier tutorials on test and interval conditions, these checks are not interchangeable.
- Treating a calculated standard error as proof that a procedure is valid. Check the procedure’s conditions explicitly. If an interval’s observed success count is too small, calculating its plug-in standard error does not repair the failed condition.
Key Takeaway
A one-proportion test and a one-proportion confidence interval use related formulas, but their standard errors serve different purposes. The test measures variation predicted by the null hypothesis. The interval estimates variation using the sample proportion. Matching the standard error and count check to the procedure keeps the calculation connected to the question being asked.
Check Your Understanding
For each question, identify which proportion belongs in the standard-error formula and explain your reasoning.
- A test has \(n=120\), \(p_0=0.40\), and \(\hat{p}=0.35\). Write the correct expression for the test standard error.
- For the same data, write the estimated standard error expression used for a one-proportion interval. Do the two standard errors have to match?
- A sample has \(n=60\), \(x=5\), and the test null value is \(p_0=0.50\). Find the test expected successes and failures, then find the interval’s observed successes and failures. Which procedure meets its Large Counts condition?
- Explain why using \(\hat{p}\) in a test standard error would not model sample-to-sample variation under \(H_0\).
- If \(\hat{p}=p_0\), what relationship holds between the test and interval standard-error calculations? Explain why that does not make the test and interval the same procedure.