Tutorials › AP Statistics › Comparing Two Independent Random Samples With a Test

Two-proportion hypothesis tests · Tutorial 533 of 1000

Comparing Two Independent Random Samples With a Test

Use a two-proportion z-test to compare independent random samples, then interpret the evidence in context without turning an association into a causal claim.

Intermediate 9 min read

What You'll Learn

  • Define the two population proportions and preserve the subtraction order throughout a comparison.
  • Match a two-sided or one-sided alternative to the research question.
  • Check random sampling, independent groups, the 10% condition, and pooled expected counts.
  • Calculate and interpret the pooled test statistic and p-value.
  • Use random sampling to guide generalization and distinguish association from causation.

Testing a Difference Between Two Surveyed Populations

A survey may ask whether two populations differ in the proportion of people who share a particular characteristic. For example, a researcher might compare the proportions of residents in two towns who use a public service. If each group is represented by its own independent random sample, a two-proportion \(z\)-test can assess whether the observed sample difference is convincing evidence of a difference between the population proportions.

This setting differs from the randomized experiments discussed in Testing a Treatment Effect in a Randomized Experiment. In a survey, people are sampled from existing populations; they are not randomly assigned to belong to either population. A test can identify evidence of an association between population membership and the characteristic, but it does not show that being in one population caused the difference.

Definition: Let \(p_1\) and \(p_2\) be the true proportions with the same defined characteristic in Population 1 and Population 2. The parameter \(p_1-p_2\) measures the difference in those population proportions, in the stated order. A test of \(H_0:p_1=p_2\) assesses whether the samples provide evidence that the proportions differ.

As explained in Hypotheses for Comparing Two Population Proportions, the alternative hypothesis must reflect the research question, not which sample proportion happens to be larger. Use \(H_a:p_1\ne p_2\) to look for any difference. Use \(H_a:p_1>p_2\) or \(H_a:p_1<p_2\) only when the question calls for evidence in that particular direction.

Conditions and What the Design Allows You to Say

For independent random samples, check how the people were selected and whether the two groups are genuinely independent. Each person should belong to only one sampled group, with no matching or pairing between observations. If sampling without replacement, check the 10% condition separately for each population: each sample size should be no more than 10% of the population size. Finally, check the Large Counts condition using the pooled proportion and the expected successes and failures in each group, as covered in Large Counts Using Expected Successes and Failures in Each Group.

Conditions: Check that each group was selected using an appropriate random sample; that the samples represent independent groups; that the 10% condition is met for each sample drawn without replacement; and that the four pooled expected counts—successes and failures in each group—are each at least 10.

Random sampling supports generalizing results to the populations from which the samples were selected, provided the sampling process is appropriate. It does not make group membership an assigned treatment. Thus, even if a test provides convincing evidence of a difference, a survey-based conclusion should describe an association between population membership and the characteristic, not a cause-and-effect relationship.

For the test calculation, the null hypothesis assumes a common population proportion. Pool the two samples’ successes and divide by their combined sample size. This pooled proportion is used to calculate the standard error and \(z\)-statistic under the null, as described in Pooled Proportion and Why It Is Used and Calculating the Pooled Standard Error. The test’s standard error is pooled; a two-proportion confidence interval instead uses separate sample proportions.

$$ \hat{p}_{\text{pool}}=\frac{x_1+x_2}{n_1+n_2}, \qquad SE_{\text{pooled}}= \sqrt{\hat{p}_{\text{pool}}(1-\hat{p}_{\text{pool}}) \left(\frac{1}{n_1}+\frac{1}{n_2}\right)}, \qquad z=\frac{\hat{p}_1-\hat{p}_2}{SE_{\text{pooled}}} $$

The p-value is the probability, assuming \(H_0\) is true, of getting a difference in sample proportions at least as extreme as the observed difference in the direction or directions specified by \(H_a\). A small p-value can be convincing evidence against equal population proportions. It is not a measure of how important a difference is, and it does not establish why the populations differ.

A Four-Step Test for Independent Random Samples

A complete response connects the study design, the calculation, and the conclusion. The order below helps keep the population claim and the test method aligned.

1
State.
Define \(p_1\) and \(p_2\) for the two populations and the shared characteristic. Write \(H_0:p_1=p_2\) and the alternative that matches the question.
2
Plan.
Identify the random samples and independent groups, check the 10% condition for each sample if applicable, and verify all four pooled expected counts.
3
Do.
Calculate the pooled proportion, pooled standard error, test statistic, and p-value for the stated alternative.
4
Conclude.
Compare the p-value with \(\alpha\), make the correct decision, and describe the evidence about the population proportions in context. Limit the conclusion to association.

Worked Examples

Worked Example: Comparing Two Towns’ Composting Proportions

Question: In a fictional survey, independent random samples of 200 households are selected from each of two towns. In Town A, 118 sampled households report regularly composting food scraps; in Town B, 94 do. Do the data provide evidence that the towns’ population proportions differ? Use \(\alpha=0.05\). Assume each town has more than 2,000 households.

State: Let \(p_1\) be the true proportion of Town A households that regularly compost food scraps, and \(p_2\) the corresponding proportion for Town B households. Test \(H_0:p_1=p_2\) against \(H_a:p_1\ne p_2\), since the question asks whether the proportions differ in either direction.

Plan: Each group was selected by an independent random sample, and no household appears in both samples, so the random sampling and independent-groups conditions are supported. Each sample of 200 is no more than 10% of its town’s population because each town has more than 2,000 households. The pooled proportion is \((118+94)/(200+200)=212/400=0.53\). Under the null model, each group has \(200(0.53)=106\) expected successes and \(200(0.47)=94\) expected failures. All four expected counts are at least 10, so the Large Counts condition is met.

Do: The sample proportions are \(\hat{p}_1=118/200=0.59\) and \(\hat{p}_2=94/200=0.47\). The observed difference is \(0.59-0.47=0.12\). The pooled standard error is:

$$ SE_{\text{pooled}} =\sqrt{0.53(0.47)\left(\frac{1}{200}+\frac{1}{200}\right)} =\sqrt{0.002491} \approx 0.04991 $$

The test statistic, using the unrounded standard error, is:

$$ z=\frac{0.59-0.47}{0.0499099} \approx 2.404 $$

For the two-sided alternative, the p-value is the probability of a standard normal result at least 2.404 units from zero in either direction: \(2P(Z\ge2.404)\approx0.0162\), rounded to four decimal places. A TI-84 2-PropZTest with \(x_1=118,n_1=200,x_2=94,n_2=200\), and the not-equal alternative gives results that round to these values.

Conclude: If the true composting proportions in the two towns were equal, the probability of obtaining a sample difference at least as far from zero as the observed difference, in either direction, would be about \(0.0162\). Since \(0.0162<0.05\), we reject \(H_0\). The samples provide convincing evidence that the proportion of households that regularly compost food scraps differs between the two towns. The observed proportion is higher in Town A, and the random samples support generalizing to the two town populations. Because this was a survey and not an experiment, the result shows an association with town, not that living in one town causes composting behavior.

Worked Example: A Difference That Is Not Convincing Evidence

Question: In a fictional survey, 100 randomly selected customers from each of two independent grocery-store populations are asked whether they use the stores’ mobile checkout feature. Forty-six customers in Group 1 and 39 in Group 2 say they do. Is there convincing evidence of a difference in the population proportions? Use \(\alpha=0.05\). Each population contains more than 1,000 customers.

Let \(p_1\) and \(p_2\) be the true proportions of customers in the two populations who use mobile checkout. The hypotheses are \(H_0:p_1=p_2\) and \(H_a:p_1\ne p_2\). Both groups are represented by independent random samples, with each sampled customer belonging to one group only. Each sample is 10% or less of its population. The pooled proportion is \((46+39)/(100+100)=85/200=0.425\). The expected successes are \(100(0.425)=42.5\) in each group, and the expected failures are \(100(0.575)=57.5\) in each group. All four expected counts exceed 10.

The sample proportions are \(\hat{p}_1=46/100=0.46\) and \(\hat{p}_2=39/100=0.39\), so the observed difference is \(0.07\). The pooled standard error and test statistic are:

$$ SE_{\text{pooled}} =\sqrt{0.425(0.575)\left(\frac{1}{100}+\frac{1}{100}\right)} =\sqrt{0.0048875} \approx 0.06991 $$
$$ z=\frac{0.46-0.39}{0.0699106} \approx 1.001 $$

The two-sided p-value is \(2P(Z\ge1.001)\approx0.3167\), rounded to four decimal places. Since \(0.3167>0.05\), we fail to reject \(H_0\). These samples do not provide convincing evidence that the mobile-checkout proportions differ between the two customer populations. This does not prove that the population proportions are equal; as discussed in Small Samples and Large P-Values, a large p-value does not rule out a real difference.

Worked Example: Testing a Predicted Direction in Two Communities

Question: A fictional survey uses independent random samples of 180 adults from each of two communities. In Community A, 105 respondents say they use a public library’s digital lending service at least once a month. In Community B, 78 say they do. Before collecting responses, the researcher asks whether the monthly-use proportion is higher in Community A. Use \(\alpha=0.05\). Each community has more than 1,800 adults.

State: Let \(p_1\) be the true proportion of Community A adults who use the digital lending service at least monthly, and \(p_2\) the corresponding proportion of Community B adults. Test \(H_0:p_1=p_2\) against \(H_a:p_1>p_2\). This upper-tail alternative matches the stated prediction.

Plan: The survey takes separate random samples from the two communities, and respondents are not paired or counted in both groups, supporting random sampling and independence. Each sample of 180 is no more than 10% of a population larger than 1,800, so the 10% condition is met. The pooled proportion is \((105+78)/(180+180)=183/360\approx0.5083\). The expected successes in each group are \(180(183/360)=91.5\); the expected failures are \(180(177/360)=88.5\). Each expected count is at least 10.

Do: The sample proportions are \(\hat{p}_1=105/180\approx0.5833\) and \(\hat{p}_2=78/180\approx0.4333\). Their difference is \(0.1500\). The pooled standard error and test statistic are:

$$ SE_{\text{pooled}} =\sqrt{\frac{183}{360}\left(1-\frac{183}{360}\right) \left(\frac{1}{180}+\frac{1}{180}\right)} \approx 0.05270 $$
$$ z=\frac{105/180-78/180}{0.0526973} \approx 2.846 $$

For the upper-tail alternative, the p-value is \(P(Z\ge2.846)\approx0.0022\), rounded to four decimal places. Since \(0.0022<0.05\), we reject \(H_0\). The data provide convincing evidence that the proportion of adults who use the digital lending service at least monthly is higher in Community A than in Community B. Because the samples are random, generalization to the respective community populations is supported, subject to the survey’s quality. The survey does not establish that living in Community A causes greater use.

Common Mistakes and AP Exam Tips

  • Treating a survey comparison like an experiment: A small p-value does not make population membership a cause. For independently sampled populations, describe evidence of a difference or association.
  • Confusing the two kinds of randomness: Random sampling supports generalizing to the population sampled. Random assignment in an experiment supports a causal interpretation. A survey’s random samples do not supply random assignment.
  • Using the sample results to choose the alternative: Choose the direction from the research question before looking at which sample proportion is larger. Otherwise, a one-sided test can be selected after the fact.
  • Checking only the combined expected counts: Verify successes and failures separately for each group under the pooled null model. All four expected counts must be at least 10.
  • Using an interval’s standard error in the test: For a test of equal proportions, use the pooled proportion in the standard error. The interval procedure uses separate sample proportions.
  • Concluding that the proportions are equal after failing to reject: A large p-value means the samples do not provide convincing evidence against the null; it does not prove equality. State the result in context without claiming no difference.
  • Leaving the groups or outcome vague: Define both population proportions using the same characteristic and preserve the order \(p_1-p_2\) in the statistic and conclusion.
AP Exam Tip: Make the design visible in your conclusion. For independent random samples, a significant result can support generalizing an association to the populations sampled. Do not claim that one population characteristic caused the other.

Key Takeaway

A two-proportion \(z\)-test compares the observed difference between two independent sample proportions with what would be expected if the population proportions were equal. In survey settings, the sampling design determines whether results can be generalized, while the absence of random assignment limits conclusions to association.

Key takeaway: Define \(p_1\) and \(p_2\), choose the alternative from the research question, check the conditions, and use the pooled test calculation. Interpret the evidence in context and describe survey findings as associations, not causes.

Check Your Understanding

For each question, assume a two-proportion \(z\)-test is appropriate unless you are asked to identify a condition.

  1. Independent random samples are taken from two school-district populations. Define \(p_1\) and \(p_2\) for a comparison of the proportions who walk to school, and write hypotheses for testing whether the proportions differ.
  2. A test of \(H_0:p_1=p_2\) gives a pooled proportion of \(0.40\), with sample sizes \(n_1=50\) and \(n_2=80\). List the four expected counts that must be checked for the Large Counts condition.
  3. A two-sided test gives a p-value of \(0.03\) at \(\alpha=0.05\). What is the decision, and what should the conclusion say about the two population proportions?
  4. Two independent random samples come from separate communities, and a test finds convincing evidence of different proportions. What can random sampling support, and why does the result not establish causation?
  5. Why should a researcher decide between \(p_1>p_2\) and \(p_1\ne p_2\) before inspecting the sample proportions?