Tutorials › AP Statistics › Confidence Intervals and Two-Sided Tests

One-proportion confidence intervals · Tutorial 434 of 1000

Confidence Intervals and Two-Sided Tests

Connect a 95% confidence interval with a two-sided test at the 0.05 level, and interpret an excluded value without overstating what the comparison proves.

Intermediate 12 min read

What You'll Learn

  • Match a 95% confidence interval with a two-sided test using a 0.05 significance level.
  • State null and alternative hypotheses for a two-sided test of a population proportion.
  • Explain why an interval excluding the null value is evidence against that value.
  • Distinguish the standard error used in a one-proportion interval from the one used in a test.
  • Use a test p-value to make the test decision when the interval and test do not agree.
  • Communicate conclusions about a population proportion in context.

From an Interval Comparison to a Test

In Using an Interval to Evaluate a Claimed Proportion, you compared a claimed value with the endpoints of a confidence interval. That comparison has a useful connection to a formal hypothesis test. In particular, a 95% confidence interval is associated with a two-sided test at the 0.05 significance level.

A two-sided test asks whether the data provide evidence that a population proportion differs from a specified value, in either direction. Write the null hypothesis as \(H_0:p=p_0\), where \(p_0\) is the claimed proportion. The alternative is \(H_a:p\ne p_0\). The significance level, written \(\alpha\), is the threshold used to decide whether the test result is statistically significant.

Definition: A two-sided test of a population proportion compares \(H_0:p=p_0\) with \(H_a:p\ne p_0\). At significance level \(\alpha=0.05\), reject \(H_0\) when the test p-value is at most 0.05; otherwise, fail to reject \(H_0\).

For a 95% confidence interval and a two-sided test, the associated significance level is \(\alpha=1-0.95=0.05\). The connection gives a helpful way to interpret the interval: if \(p_0\) is excluded, the interval provides evidence against that value; if \(p_0\) is included, the interval does not provide convincing evidence against it at the corresponding level.

However, that is an approximate connection for the usual one-proportion \(z\)-interval and one-proportion \(z\)-test. As covered in Structure of a One-Proportion \(z\)-Interval, the interval uses an estimated standard error based on \(\hat p\). The test uses a null standard error based on \(p_0\). Because these standard errors differ, the interval comparison and the test decision can occasionally disagree, especially when the claimed value is very close to an endpoint.

Why the Standard Errors Differ

A one-proportion \(z\)-interval is centered at the observed sample proportion \(\hat p\). Its estimated standard error is calculated using \(\hat p\), because the population proportion \(p\) is unknown. A one-proportion \(z\)-test instead assumes \(H_0\) is true for its calculation, so it uses \(p_0\) in the standard error.

$$ \text{Interval: }\quad \hat{p}\pm z^*\sqrt{\frac{\hat{p}(1-\hat{p})}{n}} \qquad \text{Test: }\quad z=\frac{\hat{p}-p_0}{\sqrt{\frac{p_0(1-p_0)}{n}}} $$

For a 95% interval, \(z^*\approx1.96\). For the two-sided test at \(\alpha=0.05\), the p-value is the probability, assuming \(H_0\) is true, of getting a test statistic at least as far from zero as the observed statistic in either direction. Use that test p-value—not interval membership alone—to make the formal test decision.

Important distinction: Exact decision equivalence holds when a confidence interval is formed by inverting the same test. The usual one-proportion \(z\)-interval and one-proportion \(z\)-test use different standard errors, so their comparison is a useful approximate connection, not a guarantee of identical decisions.

Conditions matter for both procedures. As in earlier tutorials on the 10% condition and the Large Counts condition, check the sampling process and the relevant counts. For the test, the Large Counts condition is checked using the null proportion: \(np_0\geq10\) and \(n(1-p_0)\geq10\). For the interval, check the observed successes and failures: \(n\hat p\geq10\) and \(n(1-\hat p)\geq10\).

Worked Examples

Worked Example: The Interval and Test Both Include the Claim

A fictional city program wants to estimate the proportion of residents who use a public bike-share service. In a random sample of 200 residents, 112 say they use it. The city has 20,000 residents. Test whether the population proportion differs from 0.50 at the 0.05 level, and compare the result with a 95% confidence interval.

1
State.
Let \(p\) be the true proportion of the city’s residents who use the bike-share service. Test \(H_0:p=0.50\) against \(H_a:p\ne0.50\) at \(\alpha=0.05\).
2
Plan and check conditions.
The sample is stated to be random. The 10% condition holds because \(200\leq0.10(20{,}000)=2{,}000\). For the test, the null expected counts are \(200(0.50)=100\) successes and \(200(0.50)=100\) failures, both at least 10. For the interval, there are 112 observed successes and \(200-112=88\) observed failures, both at least 10. Use a one-proportion \(z\)-test and a 95% one-proportion \(z\)-interval.
3
Do.
The sample proportion is \(\hat p=112/200=0.56\). The test standard error under the null is \(\sqrt{(0.50)(0.50)/200}=\sqrt{0.00125}\approx0.03536\). Thus, \(z=(0.56-0.50)/0.03536\approx1.697\), giving a two-sided p-value of approximately 0.0897. For the interval, the estimated standard error is \(\sqrt{(0.56)(0.44)/200}=\sqrt{0.001232}\approx0.03510\). Its margin of error is \(1.96(0.03510)\approx0.06880\), so the 95% interval is \(0.56\pm0.06880\), or approximately \((0.4912,0.6288)\).
4
Conclude.
Because the test p-value, 0.0897, is greater than 0.05, fail to reject \(H_0\). The sample does not provide convincing evidence that the proportion of city residents who use the bike-share service differs from 0.50. The interval includes 0.50, so its comparison likewise does not point to evidence against that value.

In this example, the interval comparison and test decision agree. The interval is a range of plausible values under its method; the test p-value measures how unusual the sample result is under the specific null model.

Worked Example: Both Procedures Give Evidence Against the Claim

A fictional school district surveys a random sample of 200 families. Ninety-six report using a district tutoring service. The district has 15,000 families. Test the claim that 60% of families use the service at the 0.05 level, and compare with a 95% interval.

Let \(p\) be the true proportion of district families who use the tutoring service. The hypotheses are \(H_0:p=0.60\) and \(H_a:p\ne0.60\). The sample is random, and the 10% condition holds because \(200\leq0.10(15{,}000)=1{,}500\). Under the null, the expected success and failure counts are \(200(0.60)=120\) and \(200(0.40)=80\), meeting the Large Counts condition. The observed interval counts are 96 successes and \(200-96=104\) failures, also both at least 10.

The sample proportion is \(\hat p=96/200=0.48\). For the test, the null standard error is

$$ \sqrt{\frac{(0.60)(0.40)}{200}} =\sqrt{0.0012} \approx0.03464 $$

The test statistic is \(z=(0.48-0.60)/0.03464\approx-3.464\). Its two-sided p-value is approximately 0.0005, so reject \(H_0\) at the 0.05 level. The sample provides convincing evidence that the proportion of district families who use the tutoring service differs from 0.60.

For the interval, the estimated standard error is \(\sqrt{(0.48)(0.52)/200}=\sqrt{0.001248}\approx0.03533\). The margin of error is \(1.96(0.03533)\approx0.06924\), yielding a 95% interval of \(0.48\pm0.06924\), or approximately \((0.4108,0.5492)\). Since 0.60 is above the upper endpoint, it is excluded. Here, the interval comparison agrees with the test decision and likewise provides evidence against the claimed proportion.

When the Interval and Test Disagree

A disagreement is possible because the interval’s standard error depends on \(\hat p\), while the test’s standard error depends on \(p_0\). If the claimed proportion is near an interval endpoint, a small difference in standard errors can place it just outside the interval while leaving the test p-value just above 0.05—or the reverse. This does not mean either calculation should be altered to force agreement.

Worked Example: A Near-Boundary Difference

A fictional environmental group takes a random sample of 1,000 households from a region of 50,000 households. In the sample, 272 households report using a rain barrel. Test whether the population proportion differs from 0.30 at the 0.05 level, and construct a 95% interval.

Let \(p\) be the true proportion of households in the region that use a rain barrel. Test \(H_0:p=0.30\) against \(H_a:p\ne0.30\). The sample is random. The 10% condition holds because \(1{,}000\leq0.10(50{,}000)=5{,}000\). Under the null, expected counts are \(1{,}000(0.30)=300\) successes and \(1{,}000(0.70)=700\) failures, both at least 10. The observed counts, 272 successes and 728 failures, also meet the Large Counts condition for the interval.

The sample proportion is \(\hat p=272/1{,}000=0.272\). For the test, the null standard error is \(\sqrt{(0.30)(0.70)/1{,}000}=\sqrt{0.00021}\approx0.01449\). The test statistic is \(z=(0.272-0.30)/0.01449\approx-1.932\), giving a two-sided p-value of approximately 0.0533. Because \(0.0533>0.05\), fail to reject \(H_0\). The sample does not provide convincing evidence at the 0.05 level that the proportion differs from 0.30.

For the interval, the estimated standard error is \(\sqrt{(0.272)(0.728)/1{,}000}=\sqrt{0.000198016}\approx0.01407\). The margin of error is \(1.96(0.01407)\approx0.02758\), giving \(0.272\pm0.02758\), or approximately \((0.2444,0.2996)\). The value 0.30 is just above the upper endpoint, so it is excluded by this 95% interval.

The interval comparison and the test decision differ: the interval excludes 0.30, but the test p-value leads to failing to reject \(H_0\). The gap is very small, and the procedures use different standard errors. Report both results accurately. Do not claim that exclusion guarantees rejection in the usual one-proportion \(z\)-test.

Common Mistakes and AP Exam Tips

  • Using interval membership as the formal test decision. With the usual one-proportion \(z\)-interval and \(z\)-test, they are approximately connected, not guaranteed to match. Calculate and use the test p-value for the test decision.
  • Using \(\hat p\) in the test standard error. Under \(H_0\), use \(p_0\): \(\sqrt{p_0(1-p_0)/n}\). The interval instead uses \(\hat p\) in its estimated standard error.
  • Calling a failure to reject proof of the null. Say the sample does not provide convincing evidence against \(H_0\) at the stated level. The test does not prove \(p=p_0\).
  • Calling an excluded value impossible. An excluded value is evidence against that value under the interval method and its conditions, not proof that it cannot be true.
  • Forgetting the two-sided alternative. When the question asks whether the proportion “differs,” use \(H_a:p\ne p_0\) and a two-sided p-value, which includes outcomes at least as extreme in either direction.
  • Leaving out the context or conditions. Identify the population proportion, show the relevant random, 10%, and Large Counts checks, and state the conclusion about the population and characteristic.
AP Exam Tip: Keep the two tasks distinct. For the interval, compare \(p_0\) with both endpoints and describe the evidence that comparison provides. For the test, calculate the statistic and two-sided p-value using the null standard error, then compare the p-value with \(\alpha\). If the results differ, explain that the usual procedures use different standard errors.

Key Takeaway

A 95% confidence interval is associated with a two-sided test at the 0.05 level: an excluded null value generally suggests evidence against that value, while an included value generally suggests the interval does not provide convincing evidence against it. For the usual one-proportion \(z\)-interval and test, treat this as an approximate connection. Their standard errors differ, so use the test p-value to decide whether to reject \(H_0\).

Key takeaway: An excluded value is evidence against the value according to the interval; it is not a guarantee that the usual two-sided \(z\)-test rejects. Check conditions, use the test p-value for the test decision, and state conclusions in context.

Check Your Understanding

Use the interval-test connection and the standard-error distinction to answer each question.

  1. What significance level is associated with a 95% confidence interval and a two-sided test?
  2. A 95% interval for a population proportion is \((0.32,0.48)\). What does excluding a claim of \(p=0.50\) suggest, and what does it not prove?
  3. In a one-proportion \(z\)-test of \(H_0:p=0.40\), which proportion belongs in the null standard error: \(\hat p\) or \(p_0\)?
  4. A two-sided test has p-value 0.07 at \(\alpha=0.05\). State the correct decision and a suitable conclusion about the evidence.
  5. Why can the usual one-proportion \(z\)-interval exclude a null value even when the corresponding \(z\)-test fails to reject it?