Tutorials › AP Statistics › Test Statistic for Two Proportions

Two-proportion hypothesis tests · Tutorial 525 of 1000

Test Statistic for Two Proportions

Calculate the two-proportion z test statistic with the pooled standard error, then explain what its sign and magnitude say about the observed difference under the null model.

Intermediate 9 min read

What You'll Learn

  • Calculate the observed difference between two sample proportions in a stated group order.
  • Use the pooled standard error to calculate the two-proportion z test statistic.
  • Interpret the sign of z as the direction of the observed difference.
  • Interpret the magnitude of z as the distance from zero in standard-error units.
  • Distinguish a test statistic from a p-value and a practical difference.
  • Check the arithmetic and communicate a test statistic in context.

From Pooled Standard Error to a Test Statistic

In Calculating the Pooled Standard Error, you learned to estimate how much the difference between two sample proportions would typically vary if the population proportions were equal. The next step is to compare the difference actually observed in the samples with that estimated variability.

The result is the test statistic, \(z\). It tells us how far the observed difference \(\hat{p}_1-\hat{p}_2\) is from the null value of zero, measured in pooled standard-error units. Its sign preserves the group order: a positive value means \(\hat{p}_1\) is larger than \(\hat{p}_2\), and a negative value means it is smaller.

Definition: For a two-proportion \(z\)-test of \(H_0:p_1=p_2\), the test statistic standardizes the observed difference \(\hat{p}_1-\hat{p}_2\) by dividing it by the pooled standard error. It measures the difference’s signed distance from zero in standard-error units.

The pooled standard error is the estimate of variability under the null hypothesis that the population proportions are equal. As established in the previous tutorial, it uses the pooled proportion \(\hat{p}_c\). The observed difference, by contrast, uses the two separate sample proportions. Keep those roles distinct.

$$ z=\frac{\hat{p}_1-\hat{p}_2}{SE_{\text{pooled}}} \qquad\text{where}\qquad SE_{\text{pooled}} =\sqrt{\hat{p}_c(1-\hat{p}_c) \left(\frac{1}{n_1}+\frac{1}{n_2}\right)} $$

The numerator and denominator are both in proportion units, so \(z\) has no units. For example, if the observed difference is \(0.05\) and the pooled standard error is \(0.025\), the difference is two standard errors above zero, giving \(z=2\). The statistic does not say that the population proportions differ by two percentage points; it says the observed difference is two estimated standard errors from the null value.

Interpretation: A positive \(z\) means the observed sample proportion in Group 1 exceeds the one in Group 2. A negative \(z\) means it is lower. The absolute value \(|z|\) gives the distance from zero in pooled standard-error units.

Worked Calculation and Interpretation

Worked Example: Comparing Appointment Reminders

Question: In a fictional study, independent random samples are taken from patients at two large clinics. At Clinic 1, 78 of 120 sampled patients attend a scheduled appointment after receiving a text reminder. At Clinic 2, 60 of 100 sampled patients attend after receiving a phone reminder. Calculate and interpret the test statistic for \(H_0:p_1=p_2\), where \(p_1\) and \(p_2\) are the true attendance proportions for the two reminder groups.

State: We are comparing the true proportions of patients who attend after the two types of reminder. Group 1 is the text-reminder group, so the difference is text minus phone. The null model says \(p_1-p_2=0\).

Plan and conditions: Use the two-proportion \(z\)-test statistic with the pooled standard error. The scenario specifies independent random samples. Assume each sample is less than 10% of its clinic’s patient population, supporting the 10% condition. For the Large Counts condition under the null model, first calculate the pooled proportion:

$$ \hat{p}_c=\frac{78+60}{120+100} =\frac{138}{220} \approx 0.6273 $$

The expected successes and failures under the null model are:

$$ \begin{aligned} \text{Clinic 1 successes: }&120\left(\frac{69}{110}\right)\approx75.27, &\quad \text{failures: }&120\left(\frac{41}{110}\right)\approx44.73,\\ \text{Clinic 2 successes: }&100\left(\frac{69}{110}\right)\approx62.73, &\quad \text{failures: }&100\left(\frac{41}{110}\right)\approx37.27. \end{aligned} $$

All four expected counts are at least 10. The design and counts support using the two-proportion \(z\)-test statistic. The conditions for this procedure are developed in detail in Conditions for a Two-Proportion z-Test.

Do: The sample proportions are \(78/120=0.65\) and \(60/100=0.60\), so the observed difference is \(0.05\). The pooled standard error is:

$$ SE_{\text{pooled}} =\sqrt{\frac{69}{110}\left(\frac{41}{110}\right) \left(\frac{1}{120}+\frac{1}{100}\right)} =\sqrt{(0.627273)(0.372727)(0.018333)} \approx\sqrt{0.004286} \approx0.06547 $$

Now divide the observed difference by the pooled standard error:

$$ z=\frac{0.65-0.60}{0.06547} =\frac{0.05}{0.06547} \approx0.764 $$

Conclude about the statistic: The observed attendance proportion for the text-reminder group is about \(0.764\) pooled standard errors above the phone-reminder group’s proportion. The sign indicates that the sample proportion is higher for text reminders. The statistic alone is not a final decision about the null hypothesis; a p-value and the chosen significance level are used for that decision.

Check: The pooled standard error is about \(0.06547\), and \(0.06547(0.764)\approx0.0500\), which returns the observed difference. This second calculation confirms the scale and arithmetic.

Reading the Sign and Size

The sign and size answer different questions. The sign gives the direction of the observed difference in the stated subtraction order. The absolute value gives how far the difference is from the null value in standard-error units. Always identify the order before describing direction: \(z<0\) means Group 1’s sample proportion is lower only when the numerator is \(\hat{p}_1-\hat{p}_2\).

A value near zero means the observed difference is small relative to its estimated sampling variability. A value with a larger absolute value means the observed difference is farther from zero relative to that variability. This is one reason the statistic is useful for inference. It is not, by itself, a measure of practical importance: the same real-world difference can produce different \(z\)-values with different sample sizes.

The alternative hypothesis determines which directions count as evidence when the p-value is calculated. For a two-sided alternative, results far in either direction are relevant; for a one-sided alternative, the specified direction matters. As covered in P-Values for One-Sided Versus Two-Sided Tests, do not choose the alternative after seeing the sign of \(z\).

Worked Example: Comparing Two Garden Programs

Question: In a fictional environmental survey, 88 of 160 households in Area 1 and 91 of 140 households in Area 2 report composting food scraps. Calculate the test statistic for the difference in population proportions, using Area 1 minus Area 2, and interpret its sign and size.

The sample proportions are \(88/160=0.55\) and \(91/140=0.65\). The observed difference is therefore \(0.55-0.65=-0.10\). For a test of equal population proportions, combine the successes and sample sizes:

$$ \hat{p}_c=\frac{88+91}{160+140} =\frac{179}{300} \approx0.5967 $$

Calculate the pooled standard error using both sample sizes:

$$ SE_{\text{pooled}} =\sqrt{\frac{179}{300}\left(\frac{121}{300}\right) \left(\frac{1}{160}+\frac{1}{140}\right)} =\sqrt{(0.596667)(0.403333)(0.013393)} \approx\sqrt{0.003223} \approx0.05677 $$

Then standardize the observed difference:

$$ z=\frac{-0.10}{0.05677}\approx-1.761 $$

The negative sign means the sample proportion in Area 1 is lower than the sample proportion in Area 2. The observed difference is about \(1.761\) pooled standard errors below zero. The statistic does not reverse the group order: it is still Area 1 minus Area 2. As an arithmetic check, \((-1.761)(0.05677)\approx-0.1000\), matching the observed difference after rounding.

Why the Same Difference Can Give a Different Statistic

The test statistic depends on both the observed difference and its pooled standard error. If the observed difference stays fixed while sample sizes increase, the pooled standard error will generally decrease. Dividing by a smaller standard error produces a larger absolute \(z\)-value. This connects to Large Sample Sizes and Tiny P-Values: more data can make a fixed difference more unusual under the null model. It does not make the difference itself larger.

Worked Example: A Difference With Larger Samples

Question: In a fictional school survey, 90 of 150 students in Program 1 and 70 of 200 students in Program 2 report bringing a reusable lunch container. Calculate the test statistic for Program 1 minus Program 2.

The sample proportions are \(90/150=0.60\) and \(70/200=0.35\). Thus, the observed difference is \(0.60-0.35=0.25\). The pooled estimate is:

$$ \hat{p}_c=\frac{90+70}{150+200} =\frac{160}{350} =\frac{16}{35} \approx0.4571 $$

The pooled standard error is:

$$ SE_{\text{pooled}} =\sqrt{\frac{16}{35}\left(\frac{19}{35}\right) \left(\frac{1}{150}+\frac{1}{200}\right)} =\sqrt{(0.457143)(0.542857)(0.011667)} \approx\sqrt{0.002895} \approx0.05381 $$

Therefore,

$$ z=\frac{0.25}{0.05381}\approx4.646 $$

The positive value indicates that the sample proportion in Program 1 is higher. The observed difference is about \(4.646\) pooled standard errors above zero. As a check, \(4.646(0.05381)\approx0.2500\). For the null-model Large Counts check, the expected successes are about \(68.57\) in Program 1 and \(91.43\) in Program 2; expected failures are about \(81.43\) and \(108.57\). Each is at least 10. The statistic describes how far the observed difference is from zero relative to variability; it does not state that the practical difference is \(4.646\) or that the programs caused the observed difference.

Common Mistakes and AP Exam Tips

  • Reversing the group order: If the numerator is \(\hat{p}_1-\hat{p}_2\), a negative statistic means Group 1’s sample proportion is lower. Name the groups and preserve that order in the interpretation.
  • Calling \(z\) a difference in proportions: The difference is the numerator, such as \(-0.10\). The statistic is the difference divided by the pooled standard error and has no units.
  • Using the unpooled standard error: For a test of \(H_0:p_1=p_2\), use the pooled standard error from Calculating the Pooled Standard Error. The unpooled standard error is used for a two-proportion confidence interval.
  • Interpreting \(z\) as a p-value: The test statistic is a standardized distance. The p-value is a probability calculated from the null model and the chosen alternative; they are not interchangeable.
  • Ignoring the sign: Reporting only \(|z|\) loses the direction of the observed difference. Give the signed statistic and explain which group has the larger sample proportion.
  • Claiming practical importance from a large \(|z|\): A large absolute statistic means the observed difference is large relative to its estimated standard error. Practical importance depends on the size and context of the difference itself.
AP Exam Tip: Show the two sample proportions, their difference in the declared order, the pooled standard error, and the division that gives \(z\). Then state what the sign and absolute value mean in context. Do not make a reject-or-fail-to-reject decision from \(z\) alone; that decision uses the p-value and significance level.

Key Takeaway

The two-proportion test statistic puts the observed sample difference on a common scale: pooled standard-error units from the null value of zero. Its sign gives the direction in the chosen group order, and its absolute value gives the distance from zero. The statistic is an important part of a test, but it is not the p-value or the test conclusion.

Key takeaway: For \(H_0:p_1=p_2\), calculate \(z=(\hat{p}_1-\hat{p}_2)/SE_{\text{pooled}}\). Interpret a positive or negative sign using the stated group order, and interpret \(|z|\) as the distance from zero in pooled standard-error units.

Check Your Understanding

For each question, keep the group order visible and distinguish the test statistic from the observed difference.

  1. If \(\hat{p}_1=0.42\), \(\hat{p}_2=0.48\), and \(SE_{\text{pooled}}=0.03\), calculate \(z\) for Group 1 minus Group 2 and interpret its sign and size.
  2. A test statistic is \(z=-2.4\) when Group 2 is subtracted from Group 1 in the order \(p_1-p_2\). Which sample proportion is larger, and how far is the observed difference from zero in standard-error units?
  3. Explain why the pooled standard error is used in the denominator when testing \(H_0:p_1=p_2\).
  4. Can a large absolute \(z\)-value by itself establish that a difference is practically important? Explain.
  5. In a two-sided test, why should the alternative hypothesis be chosen before examining the sign of the observed test statistic?