Tutorials › AP Statistics › Large Counts Using Expected Successes and Failures in Each Group

Two-proportion hypothesis tests · Tutorial 527 of 1000

Large Counts Using Expected Successes and Failures in Each Group

Use the pooled proportion to calculate expected successes and failures in each group, then judge whether the Large Counts condition supports a two-proportion z-test.

Intermediate 9 min read

What You'll Learn

  • Calculate the pooled proportion under the null hypothesis of equal population proportions.
  • Find expected successes and failures separately for Group 1 and Group 2.
  • Use a count table to organize and check the four expected counts.
  • Decide whether the Large Counts condition is met, including when a count equals 10.
  • Distinguish pooled expected counts for a test from observed counts for an interval.
  • Explain why passing the Large Counts condition does not replace other test conditions.

Turn the Pooled Proportion Into Four Expected Counts

In Conditions for a Two-Proportion z-Test, you learned that the Large Counts condition for a test of \(H_0:p_1=p_2\) uses expected counts under the null model. This tutorial focuses on calculating those counts accurately and using all four—not just the total—to decide whether the condition is met.

Under the null hypothesis, both groups are modeled as having one common success proportion. The pooled proportion, \(\hat{p}_c\), estimates that common proportion from the combined samples. For a group of size \(n_i\), multiply its size by \(\hat{p}_c\) to get its expected successes; multiply by \(1-\hat{p}_c\) to get its expected failures.

Formula: For a two-proportion test of equal population proportions, calculate \(\hat{p}_c=(x_1+x_2)/(n_1+n_2)\). Then find \(n_1\hat{p}_c\), \(n_1(1-\hat{p}_c)\), \(n_2\hat{p}_c\), and \(n_2(1-\hat{p}_c)\). The Large Counts condition is met only if each of these four expected counts is at least 10.

The calculations are a way of asking: if the two population proportions really were equal, how many successes and failures would the null model expect in each group? The counts can be decimals. For instance, an expected count of 10.5 is not a claim that half a person was observed; it is a model-based average count used to check whether the normal approximation is appropriate.

A useful shortcut is to calculate each expected count directly from the pooled total. For example, Group 1’s expected successes are \(n_1(x_1+x_2)/(n_1+n_2)\). This is the same as first finding \(\hat{p}_c\) and multiplying by \(n_1\). Whichever route you use, keep enough digits in \(\hat{p}_c\) during the calculation, and round the expected counts only when reporting them.

Organize and Check the Counts

A small table makes it harder to omit a count or mix up successes and failures. Put one row for each group and one column for each outcome. For each row, the expected successes and expected failures should add to that group’s sample size. The successes in both rows should add to \(x_1+x_2\), and the failures should add to the combined number of observed failures.

GroupExpected successesExpected failuresRow total
Group 1\(n_1\hat{p}_c\)\(n_1(1-\hat{p}_c)\)\(n_1\)
Group 2\(n_2\hat{p}_c\)\(n_2(1-\hat{p}_c)\)\(n_2\)

This table also helps reveal an important feature of the check: the pooled proportion is applied to each group’s own sample size. If one group is larger, the null model expects more successes and more failures in that group, assuming the shared proportion is between zero and one.

As explained in Large Counts for Each Group in Two-Proportion Intervals, a two-proportion confidence interval checks observed successes and failures in each group. A two-proportion test of equal proportions instead checks the expected counts calculated under the pooled null model. Do not transfer the interval’s check to the test just because both procedures involve two proportions.

Worked Examples: Calculate All Four Counts

Worked Example: Expected Counts in Two App Groups

Question: In a fictional randomized experiment, 120 participants are assigned to App 1 and 150 to App 2. A specified feature is used by 78 App 1 participants and 75 App 2 participants. Do the expected counts meet the Large Counts condition for a two-proportion \(z\)-test?

State: Let \(p_1\) and \(p_2\) be the true proportions of participants who use the feature under App 1 and App 2. For a test of \(H_0:p_1=p_2\), check the four expected counts calculated with the pooled proportion.

Plan: The Large Counts condition requires at least 10 expected successes and 10 expected failures in each group. The experiment uses separate groups of participants, and participants were randomly assigned. As in Conditions for a Two-Proportion z-Test, this calculation checks the Large Counts part of the procedure; it does not replace the other design checks.

Do: Pool the observed successes and sample sizes:

$$ \hat{p}_c=\frac{78+75}{120+150} =\frac{153}{270} \approx0.5667 $$

Using the exact fraction \(153/270=17/30\), calculate the expected counts:

$$ \begin{aligned} \text{App 1 successes: }&120\left(\frac{17}{30}\right)=68, &\quad \text{failures: }&120\left(\frac{13}{30}\right)=52,\\ \text{App 2 successes: }&150\left(\frac{17}{30}\right)=85, &\quad \text{failures: }&150\left(\frac{13}{30}\right)=65. \end{aligned} $$

All four expected counts—68, 52, 85, and 65—are at least 10.

Conclude: The pooled expected counts meet the Large Counts condition for a two-proportion \(z\)-test comparing feature-use proportions under the two apps. This result does not tell us whether the apps differ; a test statistic and p-value are needed to assess evidence.

Check: The expected counts in each row add to the sample size: \(68+52=120\) and \(85+65=150\). The expected successes total \(153\), matching the pooled success count.

Worked Example: A Small Expected Success Count in One Group

Question: In a fictional randomized study, 3 of 40 units assigned to Method 1 and 37 of 160 units assigned to Method 2 have a specified outcome. The groups are separate. Does the Large Counts condition support a two-proportion \(z\)-test?

State: Let \(p_1\) and \(p_2\) be the true proportions with the outcome under Methods 1 and 2. The question is whether all four expected counts under \(H_0:p_1=p_2\) are at least 10.

Plan: Calculate \(\hat{p}_c\), then apply it to each group’s sample size to find expected successes and failures. The other conditions, including the randomized design and independence between groups, must also be considered separately.

Do: The pooled proportion is:

$$ \hat{p}_c=\frac{3+37}{40+160} =\frac{40}{200} =0.20 $$

The four expected counts are:

$$ \begin{aligned} \text{Method 1 successes: }&40(0.20)=8, &\quad \text{failures: }&40(0.80)=32,\\ \text{Method 2 successes: }&160(0.20)=32, &\quad \text{failures: }&160(0.80)=128. \end{aligned} $$

The expected success count for Method 1 is 8, which is below 10. The other three expected counts exceed 10.

Conclude: The Large Counts condition is not met, so these expected counts do not support using a two-proportion \(z\)-test. A large combined sample size of 200 does not change the fact that one of the four group-specific counts is too small.

Check: The expected successes add to \(8+32=40\), the pooled success total. The expected failures add to \(32+128=160\), the combined failure total.

Worked Example: Expected Failures Can Be the Problem

Question: In a fictional trial with two separately assigned groups, 285 of 300 units in Group 1 and 90 of 100 units in Group 2 have a specified success. Check the Large Counts condition for a two-proportion \(z\)-test.

State: Let \(p_1\) and \(p_2\) be the true success proportions in Groups 1 and 2. Check expected successes and expected failures for each group under the null hypothesis that the proportions are equal.

Plan: Pool the successes, find the null model’s common estimated success proportion, and calculate all four expected counts. The condition fails if even one expected count is below 10.

Do: The pooled estimate is:

$$ \hat{p}_c=\frac{285+90}{300+100} =\frac{375}{400} =0.9375 $$

Thus, the estimated failure proportion under the null model is \(1-0.9375=0.0625\). The expected counts are:

$$ \begin{aligned} \text{Group 1 successes: }&300(0.9375)=281.25, &\quad \text{failures: }&300(0.0625)=18.75,\\ \text{Group 2 successes: }&100(0.9375)=93.75, &\quad \text{failures: }&100(0.0625)=6.25. \end{aligned} $$

The expected failure count for Group 2 is 6.25, below 10. The other three expected counts are at least 10.

Conclude: The pooled Large Counts condition is not met because Group 2 has fewer than 10 expected failures. High expected success counts do not compensate for a small expected failure count.

Check: The expected failures total \(18.75+6.25=25\), which matches the combined observed failure count, \(15+10=25\). The group-specific expected failure count, not the combined total, determines whether this part of the condition is met.

Worked Example: A Count of Exactly 10 Meets the Condition

Question: In a fictional randomized experiment, 4 of 25 units assigned to Treatment 1 and 36 of 75 units assigned to Treatment 2 have a specified success. Are all four expected counts at least 10?

State: Let \(p_1\) and \(p_2\) be the success proportions for the two treatments. We will check the pooled Large Counts condition for testing \(H_0:p_1=p_2\).

Plan: Calculate the pooled proportion and use it to find expected successes and failures in both groups. The wording “at least 10” includes a count equal to 10.

Do: The pooled proportion is:

$$ \hat{p}_c=\frac{4+36}{25+75} =\frac{40}{100} =0.40 $$

The expected counts are:

$$ \begin{aligned} \text{Treatment 1 successes: }&25(0.40)=10, &\quad \text{failures: }&25(0.60)=15,\\ \text{Treatment 2 successes: }&75(0.40)=30, &\quad \text{failures: }&75(0.60)=45. \end{aligned} $$

All four counts are at least 10, so the Large Counts condition is met. Notice that Treatment 1 had only 4 observed successes. For this test, the Large Counts check uses the pooled expected count of 10, not that observed count of 4.

Conclude: The pooled expected counts meet the Large Counts condition for the two-proportion \(z\)-test. The test is supported by this condition; the remaining design conditions still need to be checked before proceeding.

Check: The expected successes total \(10+30=40\), and the expected failures total \(15+45=60\). These equal the pooled observed totals.

Common Mistakes and AP Exam Tips

  • Checking only the combined totals: A combined total of at least 10 successes and 10 failures does not establish the condition. Check successes and failures within each group.
  • Using the observed counts for the test: For a test of equal proportions, calculate the expected counts using \(\hat{p}_c\). Do not use \(x_i\) and \(n_i-x_i\) as the test’s Large Counts check.
  • Checking only expected successes: A high success proportion can make expected failures small. Calculate all four values, even when the successes are plentiful.
  • Rounding too early: Keep the pooled proportion as a fraction or retain several decimal places while calculating. Premature rounding can make a borderline count look slightly above or below 10.
  • Treating an expected count as an observed count: Expected counts may be decimals because they describe what the null model predicts on average. They are used for the condition, not reported as actual sample outcomes.
  • Assuming this check settles the whole decision: Meeting the Large Counts condition does not fix a problem with randomization, sampling, independence, or the 10% condition. As in the previous tutorial, every relevant condition must be addressed.
AP Exam Tip: Show the pooled proportion and write all four expected counts with group labels. Then say explicitly whether each is at least 10. A full-credit check identifies the specific count if the condition fails; “the sample is large” or “conditions are met” is not enough.

Key Takeaway

The pooled proportion represents the common success proportion assumed by \(H_0:p_1=p_2\). Applying it to each group’s sample size gives four null-model expected counts. The test’s Large Counts condition is met only when every group has at least 10 expected successes and at least 10 expected failures.

Key takeaway: Calculate \(n_1\hat{p}_c\), \(n_1(1-\hat{p}_c)\), \(n_2\hat{p}_c\), and \(n_2(1-\hat{p}_c)\). Check each count separately; one value below 10 means the Large Counts condition is not met.

Check Your Understanding

For each question, calculate or interpret the pooled expected counts for a test of equal proportions.

  1. For \(x_1=24,n_1=60,x_2=36,n_2=90\), find \(\hat{p}_c\) and all four expected counts. Is the Large Counts condition met?
  2. Why is a combined expected success total of 15 not enough, by itself, to establish the condition?
  3. A calculation gives expected counts of 12, 28, 9.5, and 40. Which count determines the decision, and is the condition met?
  4. In a two-proportion test, one group has 7 observed successes but 11 expected successes under the pooled null model. Which count is used for the test’s Large Counts check?
  5. What does an expected failure count of 10 mean for the Large Counts condition?