Tutorials › AP Statistics › Using an Interval and a Test Together

Two-proportion hypothesis tests · Tutorial 536 of 1000

Using an Interval and a Test Together

Connect the evidence from a two-proportion test with the plausible differences shown by a confidence interval, including how to handle one-sided tests.

Intermediate 10 min read

What You'll Learn

  • Explain how a confidence interval for \(p_1-p_2\) adds information to a two-proportion test.
  • Compare a two-sided test at significance level alpha with a confidence interval at the corresponding confidence level.
  • Interpret whether an interval includes zero using the defined group order.
  • Use interval endpoints to describe plausible directions and sizes of a population difference.
  • Explain why a one-sided test and a two-sided interval may lead to different-looking results.
  • Recognize why the interval and test can differ near a decision cutoff.

One Comparison, Two Useful Summaries

A two-proportion test and a confidence interval can both help answer whether two population proportions differ, but they do different jobs. The test assesses how compatible the data are with a specified null hypothesis. The interval estimates the difference \(p_1-p_2\) and gives a range of plausible values for that difference.

As in the earlier tutorials “Interpreting a Confidence Interval for \(p_1-p_2\)” and “Writing a Conclusion for a Two-Proportion Test,” keep the group order fixed. Here, \(p_1\) and \(p_2\) are the true proportions in Groups 1 and 2 for the same outcome. The difference \(p_1-p_2\) is positive when Group 1’s proportion is higher and negative when Group 2’s is higher.

Definition: A test and an interval give complementary information about a population-proportion difference. The test provides a decision about a null value, usually zero; the interval describes plausible values for the difference, including its direction and estimated size.

For a two-sided test of \(H_0:p_1=p_2\) against \(H_a:p_1\ne p_2\), the null value for \(p_1-p_2\) is zero. A confidence interval that lies entirely above or below zero points toward a difference; one that contains zero includes “no difference” among its plausible values. That connection helps you check whether the interval and test tell a similar story.

Connecting a Two-Sided Test and Interval

A common comparison is a two-sided test at \(\alpha=0.05\) with a 95% confidence interval for \(p_1-p_2\). The confidence level is \(1-\alpha\). If zero is outside the interval, the interval indicates a difference at the matching confidence level; if zero is inside, the interval does not rule out zero.

Connection: For a two-sided test of \(H_0:p_1-p_2=0\) at significance level \(\alpha\), compare the test result with a \(100(1-\alpha)\%\) confidence interval for \(p_1-p_2\). An interval excluding zero and a rejection of \(H_0\) usually give a consistent conclusion; an interval containing zero and a failure to reject usually do as well.

There is a technical detail to keep in mind. The AP Statistics two-proportion \(z\)-test uses a pooled standard error, because its null hypothesis assumes the proportions are equal. The usual confidence interval uses an unpooled standard error, based on the two separate sample proportions. Because these calculations differ, the test and interval are not guaranteed to agree exactly when their results are very close to a cutoff. Treat the interval as a useful comparison, not a replacement for the required test decision.

When the results agree, each contributes something distinct. The test tells you whether the sample evidence is convincing enough to reject the null at the chosen \(\alpha\). The interval tells you which differences remain plausible and how large those differences might be. A p-value does not give a range of plausible effects, and an interval does not replace a conclusion that refers to the stated hypotheses and significance level.

$$ \text{Two-sided test at level }\alpha \quad\longleftrightarrow\quad \text{confidence interval with confidence level }1-\alpha $$

The arrow shows the usual conceptual pairing, not a promise of identical decisions in every borderline case when the test and interval use different standard errors. In either method, preserve the group order and describe the result in context.

Worked Examples

Worked Example: The Test Rejects and the Interval Is Above Zero

Question: In a fictional study, independent random samples of 120 households are taken from each of two towns. In Town 1, 70 households report having a home garden; in Town 2, 50 do. At \(\alpha=0.05\), test whether the true proportions differ and construct a 95% confidence interval for the difference.

State: Let \(p_1\) be the true proportion of households in Town 1 that have a home garden, and let \(p_2\) be the corresponding proportion in Town 2. Test \(H_0:p_1=p_2\) against \(H_a:p_1\ne p_2\). The parameter is \(p_1-p_2\).

Plan: The data come from independent random samples, and each household is classified using the same yes-or-no outcome. Assume each sample is less than 10% of its town’s households, satisfying the 10% condition. For the test, the pooled proportion is \((70+50)/(120+120)=0.50\). Each group has expected successes \(120(0.50)=60\) and expected failures \(120(0.50)=60\), so the Large Counts condition for the test is met. For the interval, the observed success and failure counts are 70 and 50 in Town 1 and 50 and 70 in Town 2; each is at least 10. The conditions support both procedures.

Do: The sample proportions are \(\hat{p}_1=70/120=0.5833\) and \(\hat{p}_2=50/120=0.4167\), so the observed difference is \(0.1667\). For the test, the pooled standard error and statistic are:

$$ SE_{\text{pooled}} =\sqrt{0.50(0.50)\left(\frac{1}{120}+\frac{1}{120}\right)} =\sqrt{0.0041667} \approx0.06455 $$
$$ z=\frac{0.1667}{0.06455}\approx2.582 $$

The two-sided p-value is approximately \(0.0098\). Since \(0.0098<0.05\), reject \(H_0\). For the interval, use the unpooled standard error:

$$ SE_{\hat{p}_1-\hat{p}_2} =\sqrt{\frac{0.5833(0.4167)}{120}+\frac{0.4167(0.5833)}{120}} \approx0.06365 $$

Using \(z^*=1.96\), the interval is \(0.1667\pm1.96(0.06365)\), or approximately \((0.0419,\ 0.2914)\).

Conclude: The test provides convincing evidence that the true proportions of households with home gardens differ between the two towns. The 95% interval excludes zero and is entirely positive, suggesting that Town 1’s true proportion is higher. It also estimates the difference: plausible values run from about 4.2 to 29.1 percentage points, with Town 1 higher.

Worked Example: The Test Fails to Reject and the Interval Includes Zero

Question: In a fictional survey, independent random samples of 100 readers are selected from each of two online magazines. Fifty-two readers of Magazine 1 and 45 readers of Magazine 2 say they would recommend it. At \(\alpha=0.05\), test whether the true recommendation proportions differ and find a 95% confidence interval.

State: Let \(p_1\) and \(p_2\) be the true proportions of readers of Magazines 1 and 2, respectively, who would recommend their magazine. The hypotheses are \(H_0:p_1=p_2\) and \(H_a:p_1\ne p_2\).

Plan: The two groups are independent random samples and use the same recommendation outcome. Assume each sample is less than 10% of its magazine’s reader population, so the 10% condition is met. The pooled proportion is \((52+45)/200=0.485\). In each group, the null-model expected successes are \(100(0.485)=48.5\), and expected failures are \(100(0.515)=51.5\); all four counts exceed 10. For the interval, the observed success and failure counts are 52 and 48 in Magazine 1 and 45 and 55 in Magazine 2, also all at least 10.

Do: The sample proportions are \(0.52\) and \(0.45\), giving a difference of \(0.07\). The pooled standard error is \(\sqrt{0.485(0.515)(1/100+1/100)}\approx0.07068\), so \(z=0.07/0.07068\approx0.9904\). The two-sided p-value is approximately \(0.3220\). Since \(0.3220>0.05\), fail to reject \(H_0\).

For the 95% interval, the unpooled standard error is:

$$ \sqrt{\frac{0.52(0.48)}{100}+\frac{0.45(0.55)}{100}} =\sqrt{0.004971} \approx0.07051 $$

The interval is \(0.07\pm1.96(0.07051)\), or approximately \((-0.0682,\ 0.2082)\). It contains zero.

Conclude: The data do not provide convincing evidence that the true recommendation proportions differ. The interval is consistent with that result because zero is among its plausible values. It also shows that the data are compatible with Magazine 1’s proportion being lower by about 6.8 percentage points or higher by about 20.8 percentage points. Failing to reject does not prove equality.

Worked Example: A One-Sided Test and a 95% Interval Answer Different Questions

Question: In a fictional randomized experiment, 100 seedlings are assigned to a new watering schedule and 100 to the usual schedule. Sixty-three seedlings on the new schedule and 50 on the usual schedule survive a defined dry period. Researchers test at \(\alpha=0.05\) whether the new schedule increases the survival proportion. Compare that test with a 95% two-sided interval.

State: Let \(p_T\) be the true proportion of seedlings that would survive under the new schedule, and \(p_C\) the corresponding proportion under the usual schedule. Test \(H_0:p_T=p_C\) against \(H_a:p_T>p_C\). The difference is \(p_T-p_C\).

Plan: The seedlings were randomly assigned, and each contributes to only one treatment group. The pooled proportion is \((63+50)/200=0.565\). Under the null, each group has expected successes \(100(0.565)=56.5\) and failures \(100(0.435)=43.5\), all at least 10. For the interval, the observed successes and failures are 63 and 37 for the new schedule and 50 and 50 for the usual schedule; all are at least 10. Thus the Large Counts conditions are met. Random assignment supports a causal interpretation for the experimental units.

Do: The sample proportions are \(0.63\) and \(0.50\), so the observed difference is \(0.13\). The pooled standard error is \(\sqrt{0.565(0.435)(1/100+1/100)}\approx0.07011\). Therefore \(z=0.13/0.07011\approx1.854\), and the upper-tail p-value is approximately \(0.0319\). Since \(0.0319<0.05\), reject \(H_0\).

For a 95% two-sided interval, the unpooled standard error is \(\sqrt{0.63(0.37)/100+0.50(0.50)/100}\approx0.06951\). The interval is \(0.13\pm1.96(0.06951)\), or approximately \((-0.0062,\ 0.2662)\). This interval includes zero.

Conclude: The one-sided test provides convincing evidence that the new watering schedule increases the survival proportion. The 95% two-sided interval, however, still includes zero. There is no contradiction: the test asks specifically about an increase, while the two-sided interval is paired with a two-sided question and uses a different standard error. A 90% two-sided interval, with critical value about \(1.645\), is the usual confidence-level comparison for a one-sided test at \(\alpha=0.05\); here it is approximately \((0.0157,\ 0.2443)\), entirely above zero. State clearly which question each result addresses.

What the Interval Adds

The test decision is concise: reject \(H_0\) or fail to reject \(H_0\). The interval adds detail that a decision alone cannot provide. Its endpoints show how large or small the population difference might plausibly be, and its position relative to zero indicates which direction is compatible with the data.

  • An interval entirely above zero suggests \(p_1>p_2\); an interval entirely below zero suggests \(p_1<p_2\).
  • An interval containing zero does not establish that the proportions are equal. It says that zero remains one plausible value for the difference.
  • A narrow interval indicates greater precision than a wide interval, but precision alone does not establish that a difference matters in practice.
  • Always translate the endpoints into the context and preserve the subtraction order. A difference of \(0.10\) means 10 percentage points, not necessarily a 10% increase.

An interval is not a probability statement about a fixed parameter. For example, a 95% confidence level means that the method used to construct intervals would capture the true difference in about 95% of repeated samples under its conditions. Once calculated, an interval gives plausible values for the fixed population difference; it does not assign a 95% probability that the parameter lies in that particular interval.

Common Mistakes and AP Exam Tips

  • Using the interval instead of the test decision: If asked to test hypotheses, state the p-value comparison with \(\alpha\) and say “reject” or “fail to reject.” Then use the interval as additional information.
  • Calling a zero-containing interval proof of equality: Zero is plausible, but other differences may be plausible too. Say that the interval includes zero and the test did not find convincing evidence of a difference.
  • Ignoring a one-sided alternative: A test for \(p_1>p_2\) asks a directional question. A 95% two-sided interval is not its exact confidence-level partner; identify the alternative and the confidence level being compared.
  • Assuming perfect agreement near the cutoff: The usual two-proportion test and interval use pooled and unpooled standard errors, respectively. If results are close to the decision boundary, follow the test calculation for the test decision and report the interval accurately.
  • Reporting endpoints without context: Name both groups, the shared outcome, the difference order, and what the interval endpoints mean. Use percentage points when interpreting proportion differences as percentages.
  • Claiming the interval proves causation or broad generalization: As explained in “Scope of Inference Based on How Data Were Collected,” use the study design to determine what the result can support. Random assignment can support a causal conclusion; random sampling can support generalization to the population sampled.
AP Exam Tip: For a two-sided test and a corresponding interval, report the test decision and p-value comparison, then interpret the interval’s direction, endpoints, and relationship to zero. If the interval and test seem inconsistent, check the alternative, confidence level, group order, and whether the results are near the cutoff.

Key Takeaway

A two-proportion test and confidence interval are complementary. The test evaluates evidence against a null difference; the interval describes plausible values for the population difference and its direction and size. A matching two-sided interval usually reinforces the test’s message, but the usual pooled test and unpooled interval can differ near a cutoff. Interpret both in context and let the study design set the scope of the conclusion.

Key takeaway: Use the test to make the decision required by the hypotheses and \(\alpha\); use the interval to explain which differences are plausible and how large they may be. Keep the group order fixed, distinguish one-sided from two-sided questions, and do not treat a zero-containing interval as proof of equality.

Check Your Understanding

Use the stated group order and distinguish the test question from the interval’s estimate.

  1. A 95% interval for \(p_1-p_2\) is \((0.03,\ 0.18)\). What does the interval suggest about the direction of the difference, and what does it say about zero?
  2. A two-sided test at \(\alpha=0.05\) has p-value \(0.12\), and its 95% interval contains zero. State what the test and interval jointly indicate without claiming equality.
  3. Why might a two-proportion test and its corresponding interval give different decisions when the results are close to the cutoff?
  4. A test of \(H_a:p_1>p_2\) rejects at \(\alpha=0.05\), while a 95% two-sided interval includes zero. Explain why these results are not necessarily contradictory.
  5. What information can an interval provide about the population difference that a reject-or-fail-to-reject decision alone does not provide?