Tutorials › AP Statistics › Comparing Two Proportions: Setting and Notation

Two-proportion confidence intervals · Tutorial 501 of 1000

Comparing Two Proportions: Setting and Notation

Define the two population and sample proportions, identify each group’s sample size, and use a consistent order for the difference.

Intermediate 10 min read

What You'll Learn

  • Define \(p_1\) and \(p_2\) for two specified populations or groups.
  • Distinguish the population proportions from the sample proportions \(\hat{p}_1\) and \(\hat{p}_2\).
  • Identify \(n_1\) and \(n_2\) and the success counts used to calculate each sample proportion.
  • Choose and state a group order before writing a difference of proportions.
  • Match the parameter difference \(p_1-p_2\) with the sample quantities for those same groups.
  • Recognize when a comparison involves paired data rather than two independent groups.

Set Up the Two Groups Before Comparing Them

A comparison of two proportions begins with a clear description of the groups and the characteristic being counted. For example, a school might compare the proportion of students in two programs who complete a project, or a community organization might compare the proportion of residents in two areas who use a service. Before any inference, you need to know what each proportion represents and which group is listed first.

In Defining the Parameter in Hypothesis Statements, you learned to define a population proportion by naming the population and the characteristic. The same principle applies here, but you need two definitions. As in the earlier tutorials on test conclusions, the population parameters describe the groups of interest; sample results are statistics calculated from collected data.

Definition: For two independent groups, \(p_1\) and \(p_2\) are the true proportions of the specified populations or groups that have the same defined characteristic. The corresponding sample proportions are \(\hat{p}_1\) and \(\hat{p}_2\), calculated from samples of sizes \(n_1\) and \(n_2\).

The word same matters: both proportions must refer to the same success criterion. If “success” means completing a project for Group 1, it must mean completing that project for Group 2 as well. The populations may be different, but the outcome being counted should be defined consistently.

Notation for Parameters, Counts, and Samples

Use subscripts to keep the groups straight. The subscript 1 always refers to the group you named first, and the subscript 2 refers to the group you named second. For group \(i\), where \(i\) is 1 or 2, let \(n_i\) be the sample size and let \(x_i\) be the number of sampled individuals with the characteristic. Then the sample proportion is \(\hat{p}_i=x_i/n_i\).

$$ \begin{aligned} p_1&=\text{true proportion in Population 1 with the characteristic}\\ p_2&=\text{true proportion in Population 2 with the characteristic}\\ n_1&=\text{sample size from Population 1}\\ n_2&=\text{sample size from Population 2}\\ \hat{p}_1&=\frac{x_1}{n_1} \qquad \hat{p}_2=\frac{x_2}{n_2} \end{aligned} $$

The parameters \(p_1\) and \(p_2\) are fixed, usually unknown values about the populations. The statistics \(\hat{p}_1\) and \(\hat{p}_2\) are calculated from samples, so they can vary from one sample to another. The symbols are similar on purpose: the hat marks a statistic calculated from sample data, rather than the population parameter itself.

The difference of population proportions is written \(p_1-p_2\). It describes the first group’s true proportion minus the second group’s true proportion. The corresponding sample quantities use the same order: \(\hat{p}_1\) refers to Group 1 and \(\hat{p}_2\) to Group 2. This tutorial establishes the notation; the next tutorial develops the sample difference as a point estimate.

Notation rule: Define the group order in words first. Then use that same order in every symbol and subtraction: Group 1 minus Group 2 means \(p_1-p_2\), with sample quantities labeled \(\hat{p}_1\) and \(\hat{p}_2\).

The sign of a difference depends on the order. If the first group has a larger proportion than the second, \(p_1-p_2\) is positive. If the first group has a smaller proportion, it is negative. Reversing the group order reverses the sign, even though the two groups and their data have not changed. Therefore, “compare the proportions” is not enough to define a difference: say which group comes first.

Independent Groups and a Shared Success Criterion

The notation above is for two independent groups: the observations in one group do not come as matched pairs with observations in the other group. A random sample from one population and a separate random sample from another are one common design. An experiment that randomly assigns different individuals to one of two treatments can also create two groups for comparison.

By contrast, if each person is measured twice, or each person in one group is deliberately matched with a person in the other group, the data are paired. Treating paired observations as two independent groups would ignore the matching. Always identify how the data were collected before choosing the notation and procedure. The study-design ideas in Scope of Inference Based on How Data Were Collected help clarify what the groups represent.

For a two-proportion comparison, “success” means the particular outcome whose proportion is being compared. Success is just a label for the outcome of interest; it does not mean that the outcome is good or desirable. The other outcome is the absence of that characteristic. Make the criterion precise enough that a reader can decide which observations count in each group.

Worked Examples

Worked Example: Comparing Project Completion in Two Programs

Setting: A fictional school wants to compare project completion in two independent after-school programs. It takes a random sample of 80 students from Program A, and 52 completed the project. It separately samples 100 students from Program B, and 58 completed the project.

Define the groups and characteristic: Group 1 is the population of students enrolled in Program A, and Group 2 is the population of students enrolled in Program B. The shared characteristic, or success, is completing the project.

Define the parameters: Let \(p_1\) be the true proportion of all students enrolled in Program A who completed the project. Let \(p_2\) be the true proportion of all students enrolled in Program B who completed the project. These are population proportions, not percentages calculated only from the samples.

Label sample sizes and sample proportions: Program A has \(n_1=80\) sampled students and \(x_1=52\) project completers. Program B has \(n_2=100\) sampled students and \(x_2=58\) project completers. Therefore,

$$ \hat{p}_1=\frac{x_1}{n_1}=\frac{52}{80}=0.65, \qquad \hat{p}_2=\frac{x_2}{n_2}=\frac{58}{100}=0.58 $$

Set the difference notation: Since Program A was named first, the population difference is \(p_1-p_2\), the true completion proportion for Program A minus the true completion proportion for Program B. The sample proportions are labeled in that same order: \(\hat{p}_1\) belongs to A and \(\hat{p}_2\) belongs to B.

The sample proportions describe these particular samples. They do not, by themselves, tell us the exact values of \(p_1\) and \(p_2\). Keeping those roles separate is why the hats matter.

Worked Example: State the Order for a Two-Area Comparison

Setting: A fictional community group surveys independent random samples of residents in two neighborhoods about whether they use a public bike-share service at least once a week. In North Neighborhood, 45 of 150 sampled residents report weekly use. In South Neighborhood, 63 of 180 sampled residents report weekly use.

Define the groups and characteristic: Group 1 is all residents of North Neighborhood, and Group 2 is all residents of South Neighborhood. Success is using the bike-share service at least once a week. The characteristic is identical for the two populations.

Define the population proportions: Let \(p_1\) be the true proportion of North Neighborhood residents who use the service weekly, and let \(p_2\) be the true proportion of South Neighborhood residents who use it weekly.

Identify sample quantities: For North Neighborhood, \(n_1=150\), \(x_1=45\), and

$$ \hat{p}_1=\frac{45}{150}=0.30 $$

For South Neighborhood, \(n_2=180\), \(x_2=63\), and

$$ \hat{p}_2=\frac{63}{180}=0.35 $$

Write and interpret the ordered difference: The requested population difference, using North minus South, is \(p_1-p_2\). It represents the true weekly-use proportion in North Neighborhood minus the true weekly-use proportion in South Neighborhood. The sample statistics remain attached to their group labels: \(\hat{p}_1=0.30\) for North and \(\hat{p}_2=0.35\) for South.

The second sample proportion is greater than the first, so the observed ordering is consistent with a negative North-minus-South difference. If the question instead specified South minus North, the parameter would be \(p_2-p_1\), and the direction of the difference would be reversed. The group order is part of the definition, not a cosmetic choice.

Worked Example: Distinguish Independent Groups From Paired Measurements

Setting: A fictional recreation center records whether 40 randomly selected visitors used a new check-in kiosk during one visit and whether a separate random sample of 50 visitors used the kiosk during another time period. The samples contain different visitors, and each visitor contributes one response.

In the first sample, 26 visitors used the kiosk. In the second sample, 28 visitors used it. The groups are defined by time period, and the samples contain different individuals, so this setup uses two independent groups rather than matched measurements.

Define the groups and parameters: Group 1 is visitors during the first time period, and Group 2 is visitors during the second time period. Let \(p_1\) be the true proportion of visitors in the first period who use the kiosk, and let \(p_2\) be the true proportion of visitors in the second period who use it. Success is kiosk use in both groups.

Label the sample data: For the first period, \(n_1=40\) and \(x_1=26\), so

$$ \hat{p}_1=\frac{26}{40}=0.65 $$

For the second period, \(n_2=50\) and \(x_2=28\), so

$$ \hat{p}_2=\frac{28}{50}=0.56 $$

Set the order: With the first time period as Group 1, the population difference is \(p_1-p_2\), first period minus second period. The sample proportions \(\hat{p}_1\) and \(\hat{p}_2\) use the same order.

Check the design distinction: If instead the center recorded the same 40 visitors both before and after a kiosk change, the two sets of responses would be paired by visitor. They would not be two independent groups simply because there were two measurements or two columns of data. The way observations are linked determines the design.

Common Mistakes and AP Exam Tip

  • Using \(p_1\) for a sample percentage: \(p_1\) is the true population proportion for Group 1. The sample statistic is \(\hat{p}_1=x_1/n_1\). A full-credit response defines the population parameter in context and uses the hat for the sample proportion.
  • Leaving the groups undefined: “\(p_1\) is the first proportion” does not identify a population or characteristic. State who is in Group 1 and what counts as success; then do the same for Group 2.
  • Changing the order midway: If \(p_1-p_2\) is defined as Program A minus Program B, do not describe it later as B minus A. Write the group labels beside the subscripts as you work.
  • Using different success criteria: A comparison is unclear if “success” means one outcome in Group 1 and a different outcome in Group 2. Define one consistent characteristic that can be counted in both groups.
  • Assuming equal sample sizes: The group sizes can differ. Use \(n_1\) for the first group and \(n_2\) for the second; there is no requirement that \(n_1=n_2\).
  • Calling paired data independent: Two measurements on the same individuals are linked. Check whether observations are matched before treating them as two independent groups.
  • Describing a sample statistic as the population truth: The values of \(\hat{p}_1\) and \(\hat{p}_2\) come from the samples. They are not automatically the exact population proportions \(p_1\) and \(p_2\).
AP Exam Tip: Begin by writing “Let \(p_1\) be the true proportion of [Group 1] who [meet the criterion], and let \(p_2\) be the true proportion of [Group 2] who [meet the same criterion].” Then attach \(n_1\), \(n_2\), \(\hat{p}_1\), and \(\hat{p}_2\) to those same groups and keep the chosen subtraction order throughout.

Key Takeaway

A two-proportion comparison needs two clearly defined populations, one shared characteristic, and a consistent group order. Population proportions use \(p_1\) and \(p_2\); sample sizes use \(n_1\) and \(n_2\); and sample proportions use \(\hat{p}_1=x_1/n_1\) and \(\hat{p}_2=x_2/n_2\). The notation \(p_1-p_2\) always means the first named group minus the second.

Key takeaway: Define both groups and the same success criterion, label each sample quantity with its group, and state the subtraction order explicitly. The hats distinguish sample statistics from population parameters, while the subscripts keep the two groups from being mixed up.

Check Your Understanding

For each question, use precise group labels and distinguish population parameters from sample statistics.

  1. A random sample of 90 first-year students includes 54 who commute, and a separate sample of 110 second-year students includes 55 who commute. Define \(p_1\), \(p_2\), \(n_1\), \(n_2\), \(\hat{p}_1\), and \(\hat{p}_2\) if first-year students are Group 1.
  2. In a study, Group 1 is a population of garden plots using Method A and Group 2 is a population using Method B. Write a contextual definition of \(p_1-p_2\) when success means producing at least 5 kilograms of tomatoes.
  3. A researcher defines \(p_1-p_2\) as the proportion in East District minus the proportion in West District. What would change if the research question instead specified West District minus East District?
  4. A team measures the same 30 runners before and after a training program. Should those measurements be labeled as two independent groups for a two-proportion comparison? Explain the design issue.
  5. Explain why \(\hat{p}_1\) and \(p_1\) are not interchangeable, even when the sample proportion is used to learn about the population proportion.