Why Sample Size Changes the Comparison
When groups have different numbers of observations, comparing their raw counts can give the wrong impression of what is more common within each group. A larger group may have more observations in a category simply because it has more observations overall. Relative frequency puts each count in relation to its own group’s total, so the comparison focuses on the share of each group in that category.
The previous tutorial, Common Errors When Comparing Distributions, introduced this idea using commute-time intervals. Here, we build on it by practicing how to choose the right denominator, organize relative frequencies in tables and graphs, and write a precise comparison. The general comparison habits from earlier tutorials still apply: name both groups, identify the feature, and support your statement with evidence.
The denominator is important: use the total for the group whose share you are describing. A proportion such as \(0.36\) is equivalent to \(36\%\). If every observation in a group belongs to exactly one category, that group’s relative frequencies add to \(1\), or \(100\%\), apart from small rounding differences.
When groups have different sample sizes, compare the relative frequencies for the same category or interval, not just their counts. If the sample sizes are equal, the raw counts and relative frequencies will give the same ordering of categories. Even then, proportions can make the comparison easier to interpret.
Worked Example: Comparing Responses From Two Groups
Worked Example: Comparing Responses From Two Groups
A fictional community program asks participants whether they would attend a weekend workshop. It surveys 50 people who received a printed invitation and 120 people who received an app notification. The responses are recorded below.
| Response | Printed invitation (50 people) | App notification (120 people) |
|---|---|---|
| Would attend | 18 | 30 |
| Would not attend | 32 | 90 |
| Total | 50 | 120 |
Explain why counts are not enough. More people in the app group said they would attend: 30 compared with 18. But the app group was also larger, with 120 people compared with 50. The count difference alone does not tell us which group had a greater share willing to attend.
Calculate each group’s relative frequencies. For the printed-invitation group, the proportion who would attend is \(18/50=0.36\), or \(36\%\). The proportion who would not attend is \(32/50=0.64\), or \(64\%\). For the app-notification group, the proportions are \(30/120=0.25\), or \(25\%\), who would attend and \(90/120=0.75\), or \(75\%\), who would not attend. The percentages in each group total \(100\%\).
Compare the groups in context. Although more app-notification recipients said they would attend in raw numbers, the percentage who said they would attend is greater among printed-invitation recipients: \(36\%\) compared with \(25\%\). The difference is \(36-25=11\) percentage points. In this survey, the share saying they would attend is 11 percentage points higher in the printed-invitation group.
A percentage-point difference is found by subtracting two percentages. It is not the same as saying the percentage is “11 percent higher.” Here, the statement supported directly by the subtraction is that the two reported percentages differ by 11 percentage points.
This example illustrates why the denominator should travel with the count in your working: 18 is divided by 50, while 30 is divided by 120. Dividing both counts by the same number would not calculate the within-group shares. The conclusion describes the surveyed groups; by itself, it does not establish that the invitation type caused any difference.
Comparing Relative-Frequency Histograms
Relative frequency is useful for quantitative data grouped into intervals, too. A histogram of counts shows the number of observations in each interval. When sample sizes differ, a histogram of relative frequencies instead shows the fraction of each group in each interval. To make the comparison meaningful, use the same interval boundaries for both groups.
If a group’s histogram is made from relative frequencies, the bars represent proportions or percentages rather than counts. The total relative frequency across all intervals for that group should be \(1\), or \(100\%\), provided the intervals cover all observations. This makes it possible to compare how each distribution is allocated across the same intervals, even when one group has many more observations.
Worked Example: Comparing Appointment Wait Times
A fictional clinic records appointment wait times for 40 patients who arrived in the morning and 100 patients who arrived in the afternoon. Wait times are grouped into the same intervals, measured in minutes.
| Wait time | Morning (40 patients) | Afternoon (100 patients) |
|---|---|---|
| Less than 5 minutes | 12 | 20 |
| 5 to less than 10 minutes | 18 | 50 |
| 10 to less than 15 minutes | 10 | 30 |
| Total | 40 | 100 |
Plan the comparison. The afternoon group has more patients, so the counts are not directly comparable as measures of how common each wait interval is. Calculate each interval’s relative frequency using its own group total. The matching intervals allow a fair comparison of the shares.
Calculate the morning relative frequencies. The shares are \(12/40=0.30\), or \(30\%\), under 5 minutes; \(18/40=0.45\), or \(45\%\), from 5 to less than 10 minutes; and \(10/40=0.25\), or \(25\%\), from 10 to less than 15 minutes. These total \(30\%+45\%+25\%=100\%\).
Calculate the afternoon relative frequencies. The shares are \(20/100=0.20\), or \(20\%\), under 5 minutes; \(50/100=0.50\), or \(50\%\), from 5 to less than 10 minutes; and \(30/100=0.30\), or \(30\%\), from 10 to less than 15 minutes. These total \(20\%+50\%+30\%=100\%\).
Compare in context. The largest share of patients in each group waited from 5 to less than 10 minutes. That interval includes \(45\%\) of the morning patients and \(50\%\) of the afternoon patients, so its relative frequency is 5 percentage points higher in the afternoon group. The morning group has a greater share waiting less than 5 minutes: \(30\%\) compared with \(20\%\). The afternoon group has a greater share waiting from 10 to less than 15 minutes: \(30\%\) compared with \(25\%\).
Connect the table to a graph. A relative-frequency histogram would show these percentages on the vertical axis, with matching wait-time intervals along the horizontal axis. Each group can be displayed using the same scales and intervals. Because these wait-time intervals have equal widths, comparing the heights of the corresponding bars compares percentages of the groups, rather than the larger raw counts that result from having more afternoon patients. With unequal-width intervals, use relative-frequency density for the bar heights so that bar areas represent the percentages.
A relative-frequency histogram does not make the sample sizes equal, nor does it make the observations identical. It changes the scale so that each group is represented by its within-group shares. Read the graph as a comparison of distributions, not as a count of how many patients were observed.
Choosing the Denominator in a Two-Way Table
A two-way table organizes counts for two categorical variables. It can answer different questions, and those questions may require different denominators. Ask, “A share of which group?” before dividing. If the question asks what percentage within Group A chose a response, divide by Group A’s total. If it asks what percentage of everyone who chose that response belongs to Group A, divide by the total who chose that response.
Worked Example: Identifying the Right Group Total
A fictional library asks families in two neighborhoods which time they prefer for a community event. The table shows the number of families giving each preference.
| Preferred time | Neighborhood East | Neighborhood West | Total |
|---|---|---|---|
| Morning | 36 | 42 | 78 |
| Evening | 44 | 78 | 122 |
| Total | 80 | 120 | 200 |
Question 1: Within each neighborhood, what share prefers morning? The phrase “within each neighborhood” tells us to use each neighborhood’s own total. In East, \(36/80=0.45\), or \(45\%\), prefer morning. In West, \(42/120=0.35\), or \(35\%\), prefer morning. Thus, the share preferring morning is 10 percentage points higher in Neighborhood East than in Neighborhood West.
Question 2: Among families preferring morning, what share is from East? This question uses a different denominator: all 78 families who prefer morning. The East share is \(36/78\approx0.4615\), or about \(46.2\%\). This is not the same as the \(45\%\) of East families who prefer morning. The first percentage describes a share within East; the second describes a share within the morning-preferring group.
Check the interpretation. For the comparison between neighborhoods, the relevant figures are \(45\%\) and \(35\%\), because the question asks about preferences within each neighborhood. Dividing both morning counts by 200 would answer neither within-neighborhood question. State the denominator group clearly so the reader knows what each percentage represents.
The same denominator check applies to tables of counts, bar charts, and survey summaries. A percentage without a clear reference group can be ambiguous. Phrases such as “of the East families,” “within the afternoon group,” or “among those who chose morning” identify the group used in the denominator.
A Practical Process for Unequal Group Sizes
Use this process whenever a question asks you to compare how common categories or intervals are in groups of different sizes. It makes the denominator choice explicit before you interpret the results.
Record the sample size for each group. Do not treat the largest group’s counts as directly comparable to the smaller group’s counts.
Compare the same response category or the same interval boundaries across groups.
For a within-group comparison, divide each count by the total for that group. For another question, such as the makeup of one response category, use the total for that category.
Convert proportions to percentages if useful. For a full set of categories within a group, check that the proportions sum to 1, or the percentages to 100%, allowing for rounding.
State which group has the greater share, give the percentages as evidence, and use percentage points when subtracting percentages.
Common Mistakes and AP Exam Tips
- Comparing counts as if the totals were equal. A larger count in the larger sample does not necessarily mean a larger share. Calculate each group’s count divided by its own total.
- Using the wrong denominator. “Within Group A” means divide by Group A’s total. “Among people who chose option X” means divide by the total who chose option X. Use the question’s wording to identify the reference group.
- Comparing different categories or intervals. Compare the same response or matching interval in each group. Otherwise, the numbers do not describe the same feature.
- Reporting a percentage without naming its group. “Forty-five percent prefer morning” is incomplete if it is unclear whether that is 45% of one neighborhood or 45% of all families. Identify the group in the denominator.
- Calling a percentage-point difference a percent difference. Subtracting \(45\%-35\%\) gives a difference of 10 percentage points. Make the unit of the comparison clear.
- Forgetting to check the totals. If categories cover all observations, the relative frequencies within each group should add to about 1, or 100%. A mismatch may signal an arithmetic error, a missing category, or rounding.
For full credit, show the relative-frequency calculation or report the percentages clearly, identify the group each percentage describes, and make an explicit comparison in context. “Group East had 36 morning responses and Group West had 42” gives counts, but it does not answer which group had the larger within-group share. “Forty-five percent of East families and 35% of West families prefer morning, so the share is 10 percentage points higher in East” directly answers that question.
Check Your Understanding
For each question, decide which group total belongs in the denominator before interpreting the relative frequencies.
- In Group A, 24 of 60 people choose option R. In Group B, 30 of 100 people choose option R. Calculate each group’s relative frequency for option R and state which group has the greater share.
- A survey records 15 responses in a category from a group of 30 and 24 responses in the same category from a group of 80. Why is it not enough to say that the second group has more responses in that category?
- A relative-frequency table has percentages of 25%, 50%, and 25% for three intervals in one group. What should the percentages add to, and what does that check help you notice?
- In a two-way table, 20 of the 50 people in Group North choose a morning event time. What denominator would you use to find the percentage of North participants who prefer morning?
- One group has 40% choosing a response and another has 32%. State the difference in percentage points.