Compare Probabilities Within Groups
In False Positives and Medical Testing, you used a condition to identify the reference group for a probability. The same idea helps compare groups: calculate the probability of an outcome separately within each group, then compare those conditional probabilities. This can reveal whether two categorical variables appear to be associated in a set of data.
For example, instead of asking what proportion of all students arrive late, we might compare the proportion arriving late among students who take the bus with the proportion among students who walk. The groups named after “among” determine the denominators. A percentage based on all students does not answer either within-group question.
A two-way table organizes counts according to two categorical variables. Suppose the groups are shown in rows and the outcome categories in columns. To calculate \(P(A\mid G_1)\), divide the count in the \(G_1\) row and the \(A\) column by the total count in the \(G_1\) row. Repeat within the \(G_2\) row. The row totals—not the grand total—are the denominators.
The comparison can be described with words, or by subtracting one conditional probability from the other. The difference is a difference in proportions. When converting it to a percent difference, use percentage points: for example, a change from 20% to 35% is a difference of 15 percentage points. That is not the same as saying the percentage increased by 15%.
If the groups are shown in columns instead, the same reasoning applies, but use the column totals. The important question is not whether the relevant cells happen to be in a row or column. It is: Which group is named by the condition? That group’s total belongs in the denominator.
Worked Example: Late Arrivals by Commute Group
Worked Example: Late Arrivals by Commute Group
In an invented school record, 42 of 140 students who usually take the bus arrived late on a particular morning. Among 120 students who usually walk, 18 arrived late. The two-way table summarizes the counts. Compare the conditional probabilities of arriving late.
| Usual commute | Late | Not late | Total |
|---|---|---|---|
| Bus | 42 | 98 | 140 |
| Walk | 18 | 102 | 120 |
| Total | 60 | 200 | 260 |
Define: Let \(L\) mean that a student arrived late, \(B\) mean that a student usually takes the bus, and \(W\) mean that a student usually walks. We want to compare \(P(L\mid B)\) with \(P(L\mid W)\).
Plan: The condition in each probability identifies the reference group. Use the bus row total, 140, for \(P(L\mid B)\), and the walking row total, 120, for \(P(L\mid W)\). Both groups have a positive total, and “late” and “not late” account for the students in each row.
Do: The within-group proportions are:
The difference, bus minus walk, is \(0.30-0.15=0.15\), or 15 percentage points. The bus-group late-arrival proportion is also twice the walking-group proportion, since \(0.30/0.15=2\). The difference and the ratio describe the same comparison in different ways; the difference is often easiest to explain.
Conclude: In these records, 30% of students who usually take the bus arrived late, compared with 15% of students who usually walk. The late-arrival proportion was 15 percentage points higher in the bus group. This is a description of an association in these data; it does not show that taking the bus caused students to be late.
The grand total would answer a different question. For instance, \(60/260\approx0.2308\) is the proportion of all students in the table who arrived late. It is not the late-arrival proportion in either commute group. As in Conditional Probability from a Two-Way Table, the condition tells you which part of the table to use as the reference group.
Describe the Comparison Clearly
A useful comparison does more than list two decimals. Name both groups, identify the outcome, and say which group has the larger conditional probability. When it helps, state the difference in percentage points. Include the denominator or the calculation so the reader can see that each proportion is based on its own group.
The direction matters. A statement that one probability is “higher” should identify which group has the higher value. So does the order of subtraction: bus minus walk gives a positive difference of 0.15, while walk minus bus gives \(-0.15\). Neither calculation is inherently wrong, but label the order and interpret it consistently.
A conditional-probability comparison describes the distribution of an outcome across groups. In a two-way table, this is one way to assess whether the variables are associated in the observed data. When the conditional probabilities are similar, the table shows little difference for that particular outcome comparison. When they differ substantially, the data show a clearer difference. The size of a difference alone does not establish that it is statistically significant or that it would persist in a wider population.
Worked Example: Assignment Completion by Reminder Setting
Worked Example: Assignment Completion by Reminder Setting
A teacher records whether students submitted an assignment on time and whether they had a reminder notification turned on. The invented counts are shown below. Compare the on-time submission proportions for the two reminder groups.
| Reminder setting | On time | Not on time | Total |
|---|---|---|---|
| On | 84 | 56 | 140 |
| Off | 54 | 126 | 180 |
| Total | 138 | 182 | 320 |
State: Let \(T\) mean that a student submitted on time, \(R\) mean that the reminder was on, and \(R^c\) mean that the reminder was off. Compare \(P(T\mid R)\) with \(P(T\mid R^c)\).
Plan: Use the total in each reminder group as its denominator. The group totals are 140 and 180, both greater than zero. Within each group, on time and not on time are the two recorded outcome categories, so the table allows a direct comparison of their on-time proportions.
Do: The conditional probabilities are:
The difference, reminder on minus reminder off, is \(0.60-0.30=0.30\), or 30 percentage points. The calculations use each group’s own total: \(84/140\) for the reminder-on group and \(54/180\) for the reminder-off group.
Conclude: In these records, 60% of students with reminders turned on submitted on time, compared with 30% of students with reminders turned off. The on-time submission proportion was 30 percentage points higher in the reminder-on group. This observational comparison does not establish that turning on the reminder caused the difference; other differences between students or circumstances could also be relevant.
It would be incorrect to divide both on-time counts by 320 and then compare them as if those were the group-specific probabilities. Those calculations would be \(84/320\) and \(54/320\), which describe each cell as a proportion of the entire table. For the question “among students with reminders on?” the condition is the reminder-on group, so the denominator must be 140.
Changing the Condition Changes the Question
A two-way table can also answer questions that condition on the outcome rather than on the group. In the reminder example, \(P(T\mid R)\) asks for the proportion on time among students with reminders on. The reverse conditional probability \(P(R\mid T)\) asks for the proportion with reminders on among students who submitted on time. These are different questions, with different denominators.
This direction issue is familiar from Why \(P(A\mid B)\) Is Not \(P(B\mid A)\) and False Positives and Medical Testing. Comparing groups usually means fixing each group as the condition and looking at the outcome proportion within that group. If instead you condition on the outcome, you are comparing the group composition of outcome categories.
Worked Example: Event Attendance and a Newsletter
Worked Example: Event Attendance and a Newsletter
A community center records whether residents attended an open house and whether they receive its newsletter. These invented counts let us compare attendance across newsletter groups, and then see how reversing the condition changes the question.
| Newsletter status | Attended | Did not attend | Total |
|---|---|---|---|
| Receives newsletter | 90 | 60 | 150 |
| Does not receive newsletter | 40 | 110 | 150 |
| Total | 130 | 170 | 300 |
Define: Let \(A\) mean that a resident attended and \(N\) mean that the resident receives the newsletter. To compare attendance by newsletter status, find \(P(A\mid N)\) and \(P(A\mid N^c)\).
Calculate and compare: The newsletter group has 150 residents, of whom 90 attended. The group without the newsletter also has 150 residents, of whom 40 attended.
The difference is \(0.60-0.2667\approx0.3333\), or about 33.33 percentage points, rounded. In these records, attendance was more common among residents who receive the newsletter.
Reverse the condition: Among residents who attended, 90 of 130 receive the newsletter, so \(P(N\mid A)=90/130\approx0.6923\). Among those who did not attend, 60 of 170 receive it, so \(P(N\mid A^c)=60/170\approx0.3529\). These are valid conditional probabilities, but they answer a different comparison: newsletter status within attendance groups, rather than attendance within newsletter groups.
Conclude: The table shows an association between newsletter status and open-house attendance: the recorded attendance proportion is higher among newsletter recipients. The table alone does not show that receiving the newsletter caused attendance. For example, residents already interested in the center might be more likely both to receive the newsletter and to attend.
Common Mistakes and AP Exam Tips
- Using the grand total as the denominator. For \(P(A\mid G)\), use the total in group \(G\), not the total of everyone in the table.
- Comparing counts instead of proportions. A larger group may have more outcomes simply because it has more members. Calculate each outcome count as a proportion of its own group before comparing.
- Reversing the condition. \(P(A\mid G)\) and \(P(G\mid A)\) have different reference groups. Read the words after “given” or “among” to choose the denominator.
- Calling a percentage-point difference a percent increase. A difference between 60% and 30% is 30 percentage points. A relative percent change would use a different calculation and should be stated only if requested.
- Claiming causation from an association. A table can show that proportions differ in the observed groups. By itself, that difference does not establish that group membership caused the outcome.
- Giving a conclusion without context. Say which outcome proportion is higher, identify both groups, and report the comparison with units such as percentage points.
For a clear response, define the outcome and groups, show each within-group fraction, and interpret the comparison in context. Make sure your conclusion matches the conditional probabilities you calculated. If the data come from an observational setting, describe an association rather than asserting a cause.
Check Your Understanding
Use the condition to select the reference group before choosing a denominator.
- In a two-way table, 36 of 90 students in Group A and 24 of 80 students in Group B chose a particular activity. Find and compare the conditional probabilities of choosing the activity.
- Why is the grand total usually not the denominator when calculating \(P(A\mid G)\)?
- A table shows that 70% of one group and 55% of another group have a certain response. State the difference in percentage points and identify which group has the higher response proportion.
- Explain in words the difference between \(P(\text{attended}\mid\text{newsletter})\) and \(P(\text{newsletter}\mid\text{attended})\).
- A table shows an association between a reminder setting and on-time submission. Explain why the table alone does not prove that the reminder caused the difference.