Tutorials › AP Statistics › Large Counts for Each Group in Two-Proportion Intervals

Two-proportion confidence intervals · Tutorial 506 of 1000

Large Counts for Each Group in Two-Proportion Intervals

Check each group’s success and failure counts to decide whether the Large Counts condition supports a two-proportion confidence interval.

Intermediate 8 min read

What You'll Learn

  • Identify the success and failure counts in each group.
  • Check for at least 10 successes and 10 failures separately in both groups.
  • Explain why the condition concerns each group’s sample proportion.
  • Distinguish the interval check from a check based on pooled counts.
  • Decide what a failed count means for using the standard interval procedure.

Why Check Counts in Both Groups?

In Two-Sample Independence Conditions for Proportions, you learned to check the study design before using a two-proportion confidence interval. There is another condition to check: each group must have enough observed successes and failures for the Normal approximation used by the interval to be reasonable.

This is not a check of the total number of observations alone. A large sample can still have too few successes if the characteristic is rare, or too few failures if the characteristic is very common. Because a two-proportion interval estimates a difference using a sample proportion from each group, check the counts in each group separately.

Definition: For the Large Counts condition for a two-proportion confidence interval, each group must have at least 10 successes and at least 10 failures. If \(x_i\) is the number of successes and \(n_i\) is the sample size in Group \(i\), check \(x_i\geq10\) and \(n_i-x_i\geq10\), for both \(i=1\) and \(i=2\).

How to Check Large Counts

First make sure the same outcome is being counted as a success in both groups, as discussed in Comparing Two Proportions: Setting and Notation. For example, if success means “renewed a subscription,” then count renewals in each group and count all other outcomes as failures. “Failure” simply means the outcome did not meet the defined success criterion; it does not imply that an outcome was harmful or undesirable.

For each group, find the number of successes and subtract it from that group’s sample size to find the number of failures. Then compare both counts with 10. Both comparisons must pass in Group 1, and both must pass in Group 2.

$$ \begin{aligned} \text{Group 1:}&\quad x_1\geq10 \quad\text{and}\quad n_1-x_1\geq10\\ \text{Group 2:}&\quad x_2\geq10 \quad\text{and}\quad n_2-x_2\geq10 \end{aligned} $$

You can also express the checks using each group’s sample proportion: \(n_i\hat{p}_i\geq10\) and \(n_i(1-\hat{p}_i)\geq10\), where \(\hat{p}_i=x_i/n_i\). These expressions are just the success and failure counts written using the sample proportion. Counting the observations directly is often clearest.

Key distinction: For an interval, check the observed success and failure counts separately in each group. Do not combine the groups’ counts or use a pooled proportion to decide whether this condition is met.

The reason for the check is that the interval procedure uses a Normal-based model for the difference between the sample proportions. Having enough successes and failures in each group supports using that approximation for each group’s sample proportion. Meeting the count rule does not guarantee that every aspect of the model is perfect; it is the standard AP Statistics condition for proceeding with the Normal-based interval.

This condition is different from the design checks in the previous tutorial. A study can have suitable random sampling or random assignment and still fail Large Counts. Conversely, sufficient counts do not fix a problem such as paired observations being treated as independent. Check the design and the counts as separate parts of the plan.

Also keep the interval procedure distinct from a two-proportion significance test. The interval’s Large Counts check uses each sample’s own observed success and failure counts. A test can use a different count calculation because it is built around a null hypothesis. Do not carry a test’s count check over to an interval.

Worked Examples

Worked Example: Both Groups Meet the Condition

Setting: Imagine an experiment comparing two reminder messages for renewing a recreation-center membership. Members are randomly assigned to one of the two messages. In Group 1, 52 of 80 members renew; in Group 2, 39 of 75 members renew. Let success mean “renewed.” Check the Large Counts condition for a two-proportion interval.

State: Let \(p_1\) be the true proportion of members who would renew after receiving Message 1, and \(p_2\) the corresponding proportion for Message 2. The interval would estimate \(p_1-p_2\), in that group order.

Plan: The members were randomly assigned to separate groups, and each member contributes one outcome. Given the stated setup, the groups are independent; as in Two-Sample Independence Conditions for Proportions, consider whether outcomes from different members can reasonably be treated as independent. Since this is an experiment with random assignment, the 10% condition for sampling without replacement is not the relevant check. Now check Large Counts separately in both groups.

Do: Group 1 has 52 successes and \(80-52=28\) failures. Group 2 has 39 successes and \(75-39=36\) failures.

$$ \begin{aligned} \text{Group 1:}&\quad 52\geq10,\qquad 28\geq10\\ \text{Group 2:}&\quad 39\geq10,\qquad 36\geq10 \end{aligned} $$

Conclude: Both groups have at least 10 successes and at least 10 failures. The Large Counts condition is met, so this condition supports using the Normal-based two-proportion interval. The count check alone does not calculate the interval or establish that all conditions are satisfied.

Worked Example: One Group Has Too Few Successes

Setting: Imagine two independently selected random samples of households, one from each of two towns. A survey asks whether a household has installed a particular energy-monitoring device. In Town 1, 8 of 160 sampled households have the device. In Town 2, 31 of 140 sampled households have it. Let “has the device” be success. Check Large Counts for an interval comparing the town proportions.

State: Let \(p_1\) and \(p_2\) be the true proportions of households with the device in Town 1 and Town 2, respectively. The proposed interval would estimate \(p_1-p_2\).

Plan: The samples are described as random and come from separate towns, supporting the chance-based and between-group parts of the design. For the within-group check, verify the 10% condition separately for each town if sampling was without replacement; that requires knowing the town population sizes. Regardless of that check, Large Counts must be checked using each sample’s own counts.

Do: Town 1 has 8 successes and \(160-8=152\) failures. Town 2 has 31 successes and \(140-31=109\) failures.

$$ \begin{aligned} \text{Town 1:}&\quad 8<10,\qquad 152\geq10\\ \text{Town 2:}&\quad 31\geq10,\qquad 109\geq10 \end{aligned} $$

Town 1 fails the successes check, even though its sample has 160 households. Town 2 meets both count checks.

Conclude: The Large Counts condition is not met for the two-proportion interval because Town 1 has fewer than 10 successes. A large number of failures in that group and adequate counts in Town 2 do not make up for the shortfall. Do not claim that the usual Normal-based interval is justified by the stated counts.

Worked Example: Many Successes but Too Few Failures

Setting: Imagine two random samples of customers from separate online services. Each customer is classified once according to whether they enable an optional privacy setting. In Service 1, 10 of 40 sampled customers enable it. In Service 2, 42 of 50 sampled customers enable it. Check the Large Counts condition.

State: Let \(p_1\) and \(p_2\) be the proportions of customers who enable the setting for Service 1 and Service 2. The interval of interest would estimate \(p_1-p_2\).

Plan: Suppose the samples were selected independently at random from the two customer populations. The design also requires a separate 10% check for each population if sampling without replacement. For Large Counts, the key is to count both outcomes in each sample, including the customers who did not enable the setting.

Do: Service 1 has 10 successes and \(40-10=30\) failures. Service 2 has 42 successes and \(50-42=8\) failures.

$$ \begin{aligned} \text{Service 1:}&\quad 10\geq10,\qquad 30\geq10\\ \text{Service 2:}&\quad 42\geq10,\qquad 8<10 \end{aligned} $$

Service 1 meets both requirements, including the boundary value of exactly 10 successes. Service 2 has plenty of successes but only 8 failures.

Conclude: The Large Counts condition fails because Service 2 has fewer than 10 failures. The standard interval is not supported by this condition, even though both groups together have many successes and failures. The check is group-by-group and outcome-by-outcome.

Worked Example: Writing a Complete Condition Check

Setting: Imagine two random samples of garden plots from separate community gardens. In Garden 1, 26 of 64 plots meet a defined soil-moisture target. In Garden 2, 18 of 70 plots meet the same target. Assume both samples were selected without replacement from gardens with at least 10 times as many plots as were sampled. Check whether the stated facts support the Large Counts condition for a two-proportion interval.

State: Let \(p_1\) and \(p_2\) be the true proportions of plots meeting the target in Garden 1 and Garden 2, respectively. The interval would estimate \(p_1-p_2\).

Plan: The samples are random and independent, and the stated population sizes meet the 10% condition for each sample. Check successes and failures separately in both groups before proceeding.

Do: Garden 1 has 26 successes and \(64-26=38\) failures. Garden 2 has 18 successes and \(70-18=52\) failures.

$$ \begin{aligned} \text{Garden 1:}&\quad 26\geq10,\qquad 38\geq10\\ \text{Garden 2:}&\quad 18\geq10,\qquad 52\geq10 \end{aligned} $$

Conclude: Each sample has at least 10 successes and 10 failures, so the Large Counts condition is met for both groups. Together with the stated random-sampling, independence, and 10% information, these checks support proceeding to the two-proportion interval. The interval itself still needs to be constructed and interpreted.

Common Mistakes and AP Exam Tip

  • Checking only the total sample size: A large \(n_i\) does not guarantee enough successes and failures. Calculate both counts for each group.
  • Checking only successes: A group with many successes can still have too few failures, as the Service 2 example shows. Both counts must reach 10.
  • Using combined or pooled counts: Do not add counts across groups to pass a group that falls short. The interval condition applies separately to each sample.
  • Using the wrong definition of success: State the shared outcome clearly, then count that outcome and its complement in both groups.
  • Writing “Large Counts is met” without showing the checks: A full-credit response reports the success and failure counts in each group and states whether each is at least 10.
  • Assuming this is the only condition: Passing Large Counts does not establish random sampling or assignment, independence between groups, or the 10% condition. Address those design checks separately.
AP Exam Tip: Make the group-by-group structure visible: “Group 1 has \(x_1\) successes and \(n_1-x_1\) failures; Group 2 has \(x_2\) successes and \(n_2-x_2\) failures. Each count is at least 10” or identify exactly which count is below 10. Do not replace these checks with a statement about the combined sample size.

Key Takeaway

For a two-proportion confidence interval, the Large Counts condition asks whether each group separately has enough observed outcomes of both types. Count successes and failures in Group 1, repeat in Group 2, and proceed only if all four counts are at least 10. If any one count is below 10, report that the condition is not met rather than using the standard Normal-based interval as though it were justified.

Key takeaway: Check \(x_1\), \(n_1-x_1\), \(x_2\), and \(n_2-x_2\). Each must be at least 10 for the Large Counts condition for a two-proportion confidence interval.

Check Your Understanding

For each situation, identify the success and failure counts and decide whether the Large Counts condition is met for a two-proportion confidence interval.

  1. Group 1 has 14 successes in a sample of 48, and Group 2 has 21 successes in a sample of 60. Show all four counts.
  2. Group 1 has 9 successes in a sample of 500, and Group 2 has 120 successes in a sample of 300. Which count causes a problem?
  3. Group 1 has 73 successes in a sample of 80, and Group 2 has 45 successes in a sample of 70. Does having many successes guarantee that the condition is met?
  4. Why is it not appropriate to add the two groups’ successes and failures before checking the condition for an interval?
  5. A group has exactly 10 successes and 10 failures. Does it meet the stated Large Counts condition? Explain using the inequalities.